论文精选

LLM后训练对错误奖励具有鲁棒性

精选理由

论文发现LLM后训练对错误奖励有鲁棒性,便宜验证器效果不输昂贵的,对训练成本控制有启发。

研究表明,Qwen3模型在HealthBench和PRBench基准测试中,昂贵验证器并不总是优于廉价验证器。研究测试了1.7B-8B参数的Qwen3模型,发现验证器间一致性无法识别最佳训练验证器。开源Gemma验证器能产生强训练结果。

原文 · Tanishq Abraham (论文推介)

A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards

"Across the tested domains, Qwen3 trainees (1.7B–8B on HealthBench; 8B on PRBench), evaluation splits, and frontier LLM reference judges (which we call golden verifiers), higher verifier agreement does not consistently identify the best training verifier. Expensive verifiers need not outperform inexpensive ones, and open-weight Gemma verifiers produce strong training outcomes."

link: https://t.co/Qhi1uyLX7E