Meta发布MIMESIS用户模拟器
Meta做了个能模拟真实用户行为的MIMESIS,比GPT-5.5更难对付,训练出的智能体更强。
Meta Superintelligence Labs训练了MIMESIS,一个90亿参数的用户模拟器。该模型基于真实对话和13种用户行为模式训练。在行为保真度上,MIMESIS比Claude Opus 5高13.4分。使用MIMESIS训练的智能体在9个未见过的模拟器上表现优于GPT-5.5训练的智能体。
Banger paper from Meta Superintelligence Labs on user simulators for agent training.
Agent RL setups usually let an assistant LLM play the user, so the simulated user is too cooperative and too explicit.
A fixed GPT-5.5 agent finds tau-bench tasks easier with these users than with real people.
This work trains MIMESIS, a 9B user simulator, on human conversations and 13 behavior patterns observed in real users. It beats Claude Opus 5 on behavioral fidelity by 13.4 points.
Agents trained against it outperform agents trained against GPT-5.5 under all nine user simulators they never saw.
They find that adding a coaching step that turns the simulator's private reasoning into feedback adds further gains.
Paper: https://t.co/mFhOJwKPLW