论文精选

CTE-Bench:评估状态化软件模拟器的反事实追踪

CTE-Bench: Counterfactual Trace Evaluation for Stateful Software Simulators

精选理由

CTE-Bench 评估模型能否预测干预如何改变状态化服务的未来行为,揭示了当前模型在缺乏正确反馈时的局限性。

CTE-Bench-Core-v1 包含 255 个场景,覆盖 6 个确定性 Python 服务,每个模型需进行 10,200 次预测。主要评分指标是效果步骤值匹配(VM),在干预改变的 2,476 个未来调用中,四个 API 托管模型(DeepSeek V4-Flash、Kimi K2.5、Qwen3.6-35B-A3B 和 Claude Sonnet 4.6)在正确反馈下达到 54.3%-61.5% 的效果步骤 VM。隐藏这些响应后,效果步骤 VM 降至 23.2%-28.9%,基于自我生成预测的条件下为 24.8%-33.2%,最多只有 1.2% 的场景被完全预测。

原文 · arXiv: DeepSeek

CTE-Bench: Counterfactual Trace Evaluation for Stateful Software Simulators

Coding agents change running software: they patch a service's code or overwrite its stored state, and then act on their own expectation of how the service will respond afterwards. A wrong expectation may surface only several calls later. Function-level code-execution benchmarks omit persistent service state, and agent benchmarks score the actions an agent takes or the final state it reaches. We introduce CTE-Bench, which measures whether a model can predict how an intervention changes a stateful service's future behavior, without asking it to choose actions. Each scenario gives the model Python service code, the calls and responses observed before the intervention, the intervention itself (a source edit or a state overwrite), and 40 fixed future calls; the model predicts every future response, and predictions are checked by executing the service. Three memory protocols control whether the model sees the correct earlier responses, none of them, or its own earlier predictions. CTE-Bench-Core-v1 contains 255 scenarios over six deterministic Python services, giving 10,200 predictions per model. The main score is effect-step value match (VM): exact response equality on the 2,476 future calls whose response the intervention changes. With correct earlier responses revealed, four API-hosted models (DeepSeek V4-Flash, Kimi K2.5, Qwen3.6-35B-A3B, and Claude Sonnet 4.6) reach 54.3%-61.5% effect-step VM. Hiding those responses lowers effect-step VM to 23.2%-28.9%; conditioning on self-generated predictions gives 24.8%-33.2%, and at most 1.2% of scenarios are predicted exactly end to end. Current models thus track intervention effects mainly when correct feedback is supplied, and their errors compound over a rollout. We release CTE-Bench-Core-v1 with its executable oracle, evaluation scripts, and an evaluation card mapping each claim to its protocol.