WorkGenesis:构建训练AI代理的工作场景
WorkGenesis: Building the Worlds That Teach Agents to Work
清华团队发布WorkGenesis框架,用2万工作单元训练出超越1.6T参数模型的AI代理。
WorkGenesis框架通过两项技术创新构建可执行的职业工作场景。该框架基于O*NET职业知识检索公共文件,合成工作请求和评分标准。实验显示,使用WorkGenesis合成的2万工作单元训练的Fx-Work-35B模型,在GDPvalAA-v2、APEX-Agents-AA和JobBench五个指标上平均得分31.00,超越1.6T参数的DeepSeek-V4-Preview。
WorkGenesis: Building the Worlds That Teach Agents to Work
The ability of Large Language Model (LLM) agents to complete daily and professional work is receiving increasing attention. Training such agents requires realistic work scenarios. Expert-authored occupational work is costly and slow to produce, while unconstrained synthesis often yields tasks with weak factual grounding or internally inconsistent requirements. To bridge this gap, we introduce WorkGenesis, a framework that constructs executable occupational work from real-world artifacts through two core technical innovations: (1) Evidence-Based Work Construction, which grounds each unit of work in real-world evidence by retrieving public files guided by O*NET occupational knowledge and synthesizing the surrounding context, companion materials, work request, and itemwise rubric around them; and (2) Execution-Guided Consistency Verification, which renders a reference deliverable inside the constructed work, attributes every unsatisfied rubric item to the agent, the task, or the rubric, and uses task and rubric defects as feedback to iteratively repair the work until it passes the audit. Experimental results demonstrate that Fx-Work-35B, trained with simple supervised fine-tuning (SFT) on only 20K units of work synthesized by WorkGenesis, achieves the highest scores among all comparable-scale baselines on the five reported metrics across GDPvalAA-v2, APEX-Agents-AA, and JobBench (31.00 versus 24.79 average score), and even surpasses frontier models such as the 1.6T DeepSeek-V4-Pro-Preview. These results show that WorkGenesis provides scalable training data for working agents.