论文

UMM-Reflection:用交错强化学习训练统一模型的反思能力

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

精选理由

一篇训练方法论文:让统一模型自己看图、自己改图,用整条轨迹的 RL 训练,GenEval 比 SFT 高了 12 分,推理时还不用 verifier。

论文提出 UMM-Reflection,把强化学习应用于统一多模态模型的完整反思轨迹,让模型诊断图像问题、生成修正版、再观察再诊断,整个循环联合训练。方法让同组轨迹共享初始图像,用组相对优势比较反思策略,且单个轨迹级优势同时更新反思文本和 flow-based 图像修正,无需外部 verifier。在 BAGEL 上,该方法比 SFT 的 GenEval 提高 12.05 分,并迁移到 WISE(+10.97)、OneIG-Bench(+3.48)和 T2I-CompBench++(+4.63),这些基准均未参与训练。

原文 · arXiv cs.AI

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.