论文多源确认精选

上海AI实验室提出SkillGym:用技能运行轨迹微调提升智能体能力

精选理由

上海AI实验室的SkillGym把技能文件变成沙箱任务再微调,Qwen3.5-35B-A3B在Terminal-Bench 2.1涨了19个点,连不带技能文件都更强了。

上海AI实验室发布SkillGym方法,把每个技能文件转成带代码检查器的沙箱任务,用通过校验的运行轨迹做微调。Qwen3.5-35B-A3B经此训练后,在Terminal-Bench 2.1上的成功率从39.33%升至58.43%,提升19.10个百分点。训练后的模型即使不带技能文件,在SkillsBench上拿到26.81%,超过带技能文件的基础模型的23.34%。再叠加技能文件时成绩可进一步提升到51.47%。

原文 · rohanpaul_ai

New Shanghai AI Laboratory paper finds that training on verified runs of human-written agent skills makes a model a better agent, even without the skill files.

Prompt-time skills depend on retrieval and instruction following, and fine-tuning on verified skill runs reduces that dependence.

Turning each skill file into a sandboxed task with a pass-or-fail checker produced training data that lifted Terminal-Bench 2.1 success by 19.10 points in Claude Code.

Skill files usually sit in the prompt, so they only help if the agent finds and follows them.

SkillGym turns each skill into a sandboxed task with a code checker, then trains on the runs that pass.

In Claude Code, Qwen3.5-35B-A3B jumped from 39.33% to 58.43% on Terminal-Bench 2.1. With no skill files, it scored 26.81% on SkillsBench, beating the base model with skills at 23.34%.

Loading the skills on top still helps, lifting it to 51.47%.