论文

医学视觉语言模型后训练审计:SFT 与 GRPO 对图像依赖能力的影响

From Reward Signal to Visual Utility: A Controlled Audit of Medical VLM Post-Training

精选理由

用 Qwen2.5-VL-3B 在 PMC-VQA 上做了严谨对照实验,发现答案准确率涨了,模型可能反而不看图了,做医疗 VLM 微调的都该看看。

研究在 Qwen2.5-VL-3B 上用 PMC-VQA 的 2000 道干净测试题做控制实验,比较语言模型 LoRA SFT、扩展多模态适配范围、标准 GRPO 和反事实证据目标。结果显示语言模型 LoRA SFT 使正确图像准确率仅提升 1.10 个百分点,但视觉收益事件减少 2.40 点、图像敏感度下降 5.60 点,配对记录中有 155 个视觉收益事件获得、203 个丢失。扩展适配范围得到的正确图像准确率反而低于语言模型 LoRA SFT,标准 GRPO 产生混合奖励组且干净测试准确率变化不确定。按贪心生成路径计分时,证据目标在训练集上有提升,但在相同训练剂量下对标准 GRPO 的优势在验证集上不一致。

原文 · arXiv cs.AI

From Reward Signal to Visual Utility: A Controlled Audit of Medical VLM Post-Training

Medical vision-language model (VLM) post-training is commonly evaluated through answer accuracy. We examine how changes in accuracy and training objectives relate to image-conditioned decisions in a controlled Qwen2.5-VL-3B study on PMC-VQA. We compare supervised fine-tuning (SFT) with low-rank adaptation (LoRA) restricted to the language model, expanded multimodal adaptation scopes, standard answer-only Group Relative Policy Optimization (GRPO), and a counterfactual evidence objective. On 2,000 clean-test questions, language model LoRA SFT changes correct-image accuracy by +1.10 percentage points (95% paired bootstrap CI:-0.85 to +3.05), while visual-benefit events decrease by 2.40 points and image sensitivity decreases by 5.60 points. Paired records reveal 155 acquired and 203 lost visual-benefit events. Broader adaptation yields lower correct-image accuracy than language-model LoRA SFT. Standard GRPO produces mixed-reward groups and parameter updates, with an uncertain clean test accuracy change. A generation audit reveals that canonical option scores can follow a different token path from generated answers. With scores taken along the greedy generation path, the evidence target improves on the training set; its gains over standard GRPO remain inconsistent on validation data at matched training doses. Sample-level analyses trace how evidence scores, decision margins, and generated answers change during post-training. This empirical and measurement audit identifies gaps between optimization activity, target acquisition, and useful held-out visual behavior.