清华牛津斯坦福论文:关闭思考后LLM仍会推理
这篇论文测了DeepSeek-V4-Flash等模型:关掉思考开关,开放式问题里99.9%还是忍不住推理,硬逼它只给答案准确率掉15分。
清华大学、牛津大学与斯坦福大学的论文《Thinking Inertia》发现,即使关闭思考模式,LLM仍会继续输出推理过程。DeepSeek-V4-Flash在关闭思考后,99.9%的开放式回答仍写出推理内容。是/否问题容易直接作答,多选题居中,开放式问题则持续把模型拉回推理。在开放式任务上强制只给答案,5个模型的遵从率约40%,但准确率下降约15个百分点。
New Tsinghua, Oxford, and Stanford paper finds that LLMs keep reasoning out loud even with thinking turned off, especially on open-ended questions.
With thinking disabled, DeepSeek-V4-Flash still wrote out reasoning in 99.9% of open-ended answers. A missing think tag or a short reply does not prove the model skipped reasoning.
Yes/no questions are easy to answer directly, multiple-choice questions sit in the middle, and open-ended ones keep pulling the model back into reasoning.
Forcing answer-only replies on open-ended tasks raised compliance to about 40% across 5 models but cut accuracy by about 15 points.
– arxiv. org/abs/2610.11765
Title: "Thinking Inertia: LLMs Keep Thinking When Told Not To"
- IT之家10-09 00:49原文