论文

TGDT:用预测误差筛选可信上下文,改进 Decision Transformer 长序列表现

Trust Guided Decision Transformer

精选理由

一篇改进 Decision Transformer 的新论文:发现长序列跑崩时预测误差会飙升,先筛掉不可信上下文再选动作,D4RL 上验证有效,做离线强化学习的可以看看。

arXiv 论文指出 Decision Transformer 在长 rollout 中因 conditioning 上下文漂移出训练分布而性能下降。作者提出 Trust Guided Decision Transformer(TGDT),利用模型自身 next state 预测误差升高这一信号判断上下文是否失效。TGDT 每步用 split conformal prediction 在 held out 离线数据上校准误差阈值,只保留可信的上下文后缀,再用 frozen critic 从中选价值最高的动作。在 D4RL 导航与运动控制任务上,TGDT 减少了持续高误差的运行,回报优于 vanilla Decision Transformer、基于 reset 的上下文控制和仅用价值的上下文选择。

原文 · arXiv cs.LG

Trust Guided Decision Transformer

Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each step, TGDT evaluates several recent context suffixes using rolling next state prediction error, calibrated against held out offline data via split conformal prediction. It keeps only suffixes whose error stays within the calibrated threshold, then uses a frozen critic to choose the highest value action among the trusted suffixes. This reverses the order used by value only elastic selection, where the critic may choose an action generated from a context the model itself has flagged as unreliable. Experiments on D4RL navigation and locomotion tasks show that state prediction, critic guidance, and hard context reset each solve only part of the problem. TGDT reduces persistent high error runs and improves return over vanilla Decision Transformer, reset based context control, and value only context selection.