论文精选

CPUNeSy:控制模型写入提升神经符号推理可靠性

CPUNeSy: Controlling Model Writes for Reliable Neuro-Symbolic Reasoning

精选理由

CPUNeSy 通过控制模型写入符号状态,在多跳推理任务上大幅提升可靠性,最高提升35分。

CPUNeSy 是一种服务架构,通过任务定义的谓词接口和证书门控控制模型写入符号状态。在法律和形式数学的多跳测试中,确定性执行恢复了与思维链和检索基线之间的大部分差距,全池增益最高达35.0分。在 ContractNLI 上,源重新检查移除了 DeepSeek 在两票同意后存活的四分之一错误答案。在 LeanDojo Benchmark 4 上,内核限制池匹配 BM15 召回率@15(89.3%)。

原文 · arXiv: DeepSeek

CPUNeSy: Controlling Model Writes for Reliable Neuro-Symbolic Reasoning

LLMs excel at recalling statistical patterns but degrade sharply when answers must be derived, especially on multi-hop chains. Delegating derivation to deterministic symbolic executors shifts reliability to whether model-generated premises are source-supported. We introduce CPUNeSy, a serving architecture that controls model writes to symbolic state via a task-defined predicate interface and certificate gate, abstaining when grounding passes disagree. Component analysis isolates deterministic execution, restricted grounding, agreement, and source rechecking. Experiments show deterministic execution drives most accuracy recovery on derivation-heavy tasks; controlled writes mainly improve selective reliability by withholding unsupported or inconsistent answers, at a coverage cost. On multi-hop tests in law and formal math, deterministic execution recovers most of the gap over chain-of-thought and retrieval baselines, with full-pool gains up to 35.0 points. Certification is selective-serving control, not accuracy mechanism: with grounding traces fixed on ContractNLI, source rechecking removes a quarter of DeepSeek's wrong answers surviving two-vote agreement, at measurable coverage cost. When abstention is costly, routing withheld cases to an uncertified same-model fallback raises full-pool accuracy on MedCalc-Bench Verified by 13.9 and 4.9 points for Seed and DeepSeek; these gains are not from the certified channel. On LeanDojo Benchmark 4, kernel-restricted pools match BM25 recall@15 (89.3%). Gains depend on the grounder's error regime: bias-dominated grounders benefit less, consistent with our voting bound. Certificates guarantee derivational validity relative to admitted premises; semantic faithfulness to natural-language sources remains conditional on the source checker, and prospective validation is future work.