EHRAdapt:用语义先验将冻结 LLM 适配电子健康记录
EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events
一篇很实用的论文:教你怎么把病人电子病历喂给冻结的 Llama3.2 或 OLMo2,只训练不到 1% 的参数,罕见病预测还更好。
EHRAdapt 是一种适配器,把电子健康记录的(时间、模态、编码)元组直接映射进冻结语言模型的嵌入空间,避免将记录序列化为文本。为解决临床词表长尾、罕见事件样本稀缺的问题,它把事件向量拆为两部分:来自生物医学模型的冻结语义先验,加上随证据积累的低秩残差修正。团队在约 400 万患者的记录上做继续预训练,使用 OLMo2 1B、Llama3.2 1B 和 OLMo2 7B 三个骨干模型,仅训练占参数 0.1%–0.6% 的适配器。结果显示移除语义通路对罕见事件预测的损害是最常见事件的十倍以上;在法定传染病和综合征分类任务上,EHRAdapt 超过基于文本的 LLM 和计数基线。
EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events
Electronic health records (EHRs) encode clinical histories as (time, modality, code) tuples, whereas pretrained language models expect text tokens. Serializing them as text inflates sequence length and redundantly encodes structure. We introduce EHRAdapt, an adapter that maps tuples directly into a frozen language model's embedding space. Modality receives a learned embedding, time gaps enter through learned attention biases, and event codes receive dedicated vectors. Learning event vectors is the central challenge: clinical vocabularies are long-tailed, leaving rare events too few observations for reliable estimates. EHRAdapt therefore represents each event vector as the sum of a semantic prior and an evidence residual. The prior is a frozen embedding of the event's clinical description from a biomedical language model trained on clinical ontologies, mapped into the model's input space by a shared learned projection, so it supplies clinical meaning even when observations are scarce. The residual, a learned low-rank event-specific correction, refines it as evidence accumulates. We run continued pretraining on about 4 million patients' records with three frozen LLM backbones (OLMo2 1B, Llama3.2 1B, and OLMo2 7B), training only the adapter (0.1--0.6% of all parameters). The full adapter outperforms all ablations in held-out next-event prediction on every backbone. Removing the semantic pathway hurts rare events over ten times more than the most frequent ones, whereas removing the residual hurts overall prediction but improves it for the rarest events. On reportable infectious-disease and syndromic downstream classification tasks, EHRAdapt outperforms text-based LLM and count-based baselines, and both pathways improve rare-disease discrimination. The two pathways therefore play complementary roles, visible only when results are broken down by event frequency rather than averaged.