模型精选

MetaCtrl:大模型元认知控制器提升推理效率

MetaCtrl: Your Large Language Models Can Reason Better and More Concisely with a Metacognitive Controller

精选理由

MetaCtrl 让大模型在保持准确率的同时大幅减少推理步骤,无需重新训练原模型就能直接使用。

MetaCtrl 是一个轻量级控制器,可自适应调节冻结的推理模型。该模型通过强化学习训练,在数学、科学和代码等七个基准测试中,将 DeepSeek-R1-Distill-Qwen-7B 的平均准确率提高 4.7 个百分点,同时减少 53.3% 的推理长度。无需额外训练即可迁移到 Qwen3-14B 等未见过的推理模型。

原文 · arXiv: DeepSeek

MetaCtrl: Your Large Language Models Can Reason Better and More Concisely with a Metacognitive Controller

Large reasoning models improve performance on challenging problems by allocating additional computation before answering, but longer reasoning does not always lead to better results and can introduce substantial redundant reasoning on simple problems. Conversely, aggressively shortening reasoning can degrade performance on difficult ones. Effective reasoning therefore requires dynamically deciding when additional computation is useful based on the reasoner's capabilities and evolving solution state. Existing approaches often rely on predefined budgets or intervention rules, retrain the target reasoner, or require additional supervision. We introduce MetaCtrl, a lightweight controller that adaptively regulates a frozen reasoner without predefined token budgets or reasoner retraining. We formulate reasoning regulation as a sequential metacognitive control problem: MetaCtrl observes the evolving reasoning trace and decides whether to continue, simplify, skip redundant steps, or conclude reasoning. It is trained directly with reinforcement learning using a reward that prioritizes correctness while favoring shorter trajectories among correct solutions, requiring neither supervised intervention trajectories nor problem-specific budgets. Across seven benchmarks spanning mathematics, science, and code, MetaCtrl consistently improves the accuracy of LRMs while reducing their reasoning length. On DeepSeek-R1-Distill-Qwen-7B, it improves average accuracy by 4.7 points while reducing generation length by 53.3%. Without further training, the same controller transfers to an unseen reasoner (e.g., Qwen3-14B), improving average accuracy by 2.9 points and reducing generation length by 50.3%. These results establish MetaCtrl as a plug-and-play controller for improving reasoning accuracy while substantially reducing inference-time generation. The code is available at https://github.com/binbin2xs/MetaCtrl.