技巧83°

图像模型后训练奖励设计方法

精选理由

Arena团队发布了图像模型后训练新方法,通过复合奖励显著提升FLUX和Ideogram性能,解决奖励攻击问题。

研究提出复合奖励机制,结合Bradley-Terry偏好模型、忠实度奖励、约束奖励和反奖励攻击标准。该方法使FLUX.2-dev在T2I排行榜上提升69分达1202分,Ideogram 4提升20分达1224分超越所有开源模型。Gemini 3.5 Flash评估显示,添加忠实度和约束奖励后胜率提升至64.2%,集成策略后达66.0%。

原文 · lmarena.ai

How to design rewards for post-training frontier image models? Our research suggests human preference reward is necessary, but insufficient: A preference model may still reward outputs that look appealing but miss details, introduce unrequested content, or exhibit other forms of reward-hacking. We therefore optimize towards a composite reward: - Bradley-Terry reward model trained on ~5.6M pairwise human votes - Faithfulness reward from auto-generated prompt checklists evaluated by a vision-language model - Constraint reward covering explicit and implicit user intent - Anti-reward-hacking rubric rewards targeting failures such as garbled text and photorealism drift This post-training recipe improves two already-strong open image models: - Post-trained FLUX.2-dev gains 69 Elo points on our live T2I leaderboard, scoring 1202 - Post-trained Ideogram 4 gains 20 Elo points reaching a score of 1224 and surpassing all publicly listed open models (as of Sep 04, 2026). Offline ablations with Gemini 3.5 Flash as the judge, show that these reward components are complementary: win rate against the base model increases as we add faithfulness and then constraint rewards on top of preference-only training, reaching 64.2%. Finally, we ensemble policies trained with and without the anti-reward-hacking objective directly in weight space, further increasing win rate to 66.0%. 💬 9 🔄 4 ❤️ 25 👀 5206 📊 11 ⚡