论文

Gemma-3-12B 微调模型匹配 GPT-4o 放射报告实体提取

Comparison of techniques for fine-tuning open-weight models for entity extraction from radiology reports

精选理由

医学 AI 研究发现:用真实数据蒸馏微调的开源模型,能在单张消费级 GPU 上匹敌 GPT-4o 的放射报告提取能力。

研究人员使用 Gemma-3-12B 开源模型进行微调,在颅内出血严重程度提取任务中达到与 GPT-4o 相当的性能(宏 F1 值 0.845 vs 0.850)。实验采用 2x2 设计,结合两种适应策略和两种训练数据源,在 100 份专家裁定报告上进行了基准测试。蒸馏真实报告数据的模型表现最佳,而合成数据模型在所有训练规模下均未超过未微调的基础模型。

原文 · arXiv cs.LG

Comparison of techniques for fine-tuning open-weight models for entity extraction from radiology reports

Converting free-text radiology reports into structured labels supports cohort building, quality assurance, and monitoring of clinical imaging models, but the strongest label extractors are hosted proprietary models whose use raises privacy, cost, and reproducibility concerns. We asked whether a fine-tuned open-weight model (Gemma-3-12B) can match GPT-4o at multi-label intracranial hemorrhage (ICH) acuity extraction from non-contrast head-CT reports, and which ingredients matter. Using a 2x2 design, we crossed two adaptation strategies (a discriminative classification head, CH; generative instruction fine-tuning, IFT) with two training-data sources (distillation of real GPT-4o-labeled reports; synthetic reports generated by GPT-4o from real exemplars), across five training sizes, benchmarked on 100 expert-adjudicated reports against GPT-4o and the un-tuned open-weight base. The distilled instruction-tuned model (DIFT) matched GPT-4o (macro-F1 0.845 vs 0.850; p = 1.000) and exceeded the base model by 0.178. The decisive factor was the training-data source, not the fine-tuning method: both synthetic-data models failed to exceed the un-tuned open-weight base at any training size and underperformed the distilled models across all acuity classes. Fine-tuning and inference fit within the memory envelope of a single 24 GB consumer GPU. For narrow, high-value clinical label-extraction tasks, distilling real reports, rather than generating synthetic ones, is what closes the gap to a hosted model, enabling a private, low-cost, version-stable on-premises alternative.