UMM-Reflection:用交错强化学习训练统一模型的反思能力
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
一篇训练方法论文:让统一模型自己看图、自己改图,用整条轨迹的 RL 训练,GenEval 比 SFT 高了 12 分,推理时还不用 verifier。
论文提出 UMM-Reflection,把强化学习应用于统一多模态模型的完整反思轨迹,让模型诊断图像问题、生成修正版、再观察再诊断,整个循环联合训练。方法让同组轨迹共享初始图像,用组相对优势比较反思策略,且单个轨迹级优势同时更新反思文本和 flow-based 图像修正,无需外部 verifier。在 BAGEL 上,该方法比 SFT 的 GenEval 提高 12.05 分,并迁移到 WISE(+10.97)、OneIG-Bench(+3.48)和 T2I-CompBench++(+4.63),这些基准均未参与训练。
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.