论文精选73°

ReFigBench 论文:同一模型在不同编码智能体框架中成绩差异明显

Impressive paper showing how much the harness changes a coding agent's results. Harnesses do play a...

精选理由

同一模型放进 Claude Code 和 Codex 跑同样的任务,结果居然会一个升一个降,这篇论文用 PowerPoint 重建实验讲清楚了 harness 的影响。

ReFigBench 要求编码智能体把真实 arXiv 论文图表重建成可编辑的 PowerPoint 幻灯片,保留文字、布局和连接线。GPT-5.5 在相同 1,000 个任务上分别跑在 Claude Code 和 Codex 两个 harness 中,配合专门的 PowerPoint 工作流后,一个框架里成绩提升、另一个反而下降,即使提示词完全相同。该基准覆盖 GPT、Claude、MiMo、MiniMax 四个家族的十种配置,采用产物检查、两类 LLM 评审和盲测人工对比打分。结果显示感知仍是主要瓶颈:专门工作流在所有配置中都去掉了原生连接线,但人工评审在多数配对中仍更偏好其渲染结果。

原文 · elvis

Impressive paper showing how much the harness changes a coding agent's results. Harnesses do play a...

Impressive paper showing how much the harness changes a coding agent's results. Harnesses do play a huge role in what you are getting out of the models. GPT-5.5 was run inside both Claude Code and Codex on the same 1,000 tasks. With a specialized PowerPoint workflow, it improved inside one harness and got worse inside the other. The harness also changed scores when the prompt was identical. ReFigBench asks coding agents to rebuild real arXiv overview figures as editable PowerPoint slides that keep the text, layout and connections. It covers ten configurations across the GPT, Claude, MiMo and MiniMax families, scored by artifact checks, two families of LLM judges and blinded human comparisons. Perception is still the main bottleneck. The specialized workflow removed native connectors in every configuration, yet human judges still preferred its renderings in most matchups. Paper: arxiv.org/abs/2609.18844 Chat with Paper: academy.dair.ai/papers/refigbe… 💬 5 🔄 3 ❤️ 24 👀 1769 📊 8 ⚡