论文

DSSR:用读者损失训练 LLM 智能体的状态写入器

Decision-Sufficient State Representations: Measuring and Reducing Write-Time Regret

精选理由

Long-context 智能体靠摘要记事,写漏一步就全完。这篇论文把摘要器的损失量化拆解,还给出了直接用读者损失训练写入器的 DSSR 方法,做 agent 记忆的值得细看。

arXiv 论文提出 DSSR(decision-sufficient state representations),针对长任务中 LLM 智能体的状态摘要丢失信息问题。作者把损失拆为预算损失和写入时遗憾两部分:在 TextWorld 烹饪游戏中,128 token 的事实型状态几乎全胜,而提示词驱动的语言模型写入器最高只赢 17%。DSSR 用读者按写入状态行动后的表现来打分候选状态,该前滚分数与游戏结果相关性达 ρ=0.48,而惯用的固定上下文打分方式 ρ≤0.07。在预注册测试集上,训练使短延迟场景成功率提升 +7.0 [+1.9, +12.2] 个百分点,但延迟越长收益越小,瓶颈在于逐步打分无法覆盖跨步骤的信用分配。

原文 · arXiv cs.LG

Decision-Sufficient State Representations: Measuring and Reducing Write-Time Regret

Long tasks produce more history than an LLM agent can hold in its context, and more than it uses reliably even when the history fits. A growing line of work therefore has agents carry a short written state instead: at every step a writer rewrites the state, and a reader acts from the state alone. Steps stay cheap, but anything the writer drops is lost before later decisions reveal that they need it. We quantify this loss and ask whether training can reduce it. Comparing the written state with the best state of the same size written in hindsight, we split the reader's loss into a budget loss, which any state of that size must incur, and a write-time regret, which comes from the writer's choices. In TextWorld cooking games where we control how long a fact must be carried before it is needed, a 128-token state holding the facts wins nearly every game, while prompted language-model writers win at most 17%. Almost all of the loss is write-time regret, and it grows with the delay. We then train the writer from the reader's own loss. DSSR (decision-sufficient state representations) scores candidate states by how well the reader acts after the writer carries them forward, and teaches the writer to prefer the better ones. This forward-rolled score predicts game outcomes ($ρ= 0.48$), whereas scoring a candidate as a fixed context, as hindsight methods usually do, does not ($ρ\leq 0.07$). On a pre-registered test split opened once, training adds +7.0 [+1.9, +12.2] points of success when facts are needed soon, bringing a plain summary writer to the level of belief- and slot-based memory prompts. The gain shrinks as the delay grows and is significant only at the shortest delay. We trace this limit to credit assignment: keeping a fact now pays off only if every later rewrite keeps it too, which a per-step score cannot see.