Meta研究:基础模型优于强化学习微调模型
Meta发现基础模型在某些任务上反而比强化学习微调版本表现更好,提出了'锐化税'概念和PTGS解决方案。
Meta Superintelligence Labs研究发现,在足够样本下,带轻量约束的基础模型能解决更多智能体任务。在BFCL v4多轮对话、ACEBench和WebShop基准测试中,当K值较大时,基础模型经常解决微调模型无法完成的任务。研究人员将这种现象称为'锐化税',并提出了PTGS解决方案。
Banger paper from Meta Superintelligence Labs.
They find something super interesting and unexpected.
(bookmark it)
Base models with a light harness often solve more agentic tasks than their RL post-trained versions when both get enough samples.
Post-trained models win on pass@1. At large K, base models frequently solve tasks the post-trained ones never solve on BFCL v4 multi-turn, ACEBench, and WebShop.
This is because post-training pushes each task toward always solved or never solved. Consistency goes up, and coverage goes down.
The authors call the lost test-time scalability the Sharpening Tax. Across 42 base and post-trained pairs, it shows up in most settings, grows with model size, and can be estimated from a few rollouts.
Their fix, PTGS, sets the sampling temperature per prompt from its estimated difficulty during RL. It pays a smaller tax and also raises pass@1.
Paper: https://t.co/HteGrBj89J