论文

SRHarness 发布:让 LLM 做符号回归的专用运行框架

SRHarness: A Harness for Agentic Symbolic Regression

精选理由

给符号回归套了个专用运行框架,同样的 DeepSeek 底座准确率从 62% 拉到 94%,还超过了 Codex,做 AI for Science 的可以看看。

arXiv 论文 SRHarness 面向 agentic 符号回归任务,提供可组合科学操作、持久化科学状态和搜索轨迹生命周期管理三套机制。在 LLM-SRBench 上,配合 DeepSeek-v4-flash-0731,SRHarness 在 LSR-Transform 达到 93.69% 符号准确率,对比 SR-Scientist 的 62.16%。在去除科学描述的匿名变体上仍保持 72.97%,而 SR-Scientist 降至 39.64%。同一模型底座下还超过 Codex(72.97% 对 20.72%),接近 Codex 加 GPT-5.5 的水平。

原文 · arXiv: DeepSeek

SRHarness: A Harness for Agentic Symbolic Regression

Recent agentic symbolic regression approaches increasingly rely on large language models to analyze data, select scientific operations, and refine hypotheses over long search trajectories. In such systems, performance depends not only on the underlying model and search strategy, but also on the runtime infrastructure that supports scientific search. We introduce SRHarness, a domain-specific harness for agentic symbolic regression built around three mechanisms: composable scientific actions that provide a common interface over raw, transformed, and candidate-derived quantities; persistent scientific state that retains evaluated hypotheses and exposes compact model-facing views; and trajectory lifecycle management that coordinates continuation, branching, restart, and termination. On LLM-SRBench, SRHarness consistently improves both numerical generalization and symbolic recovery under matched LLM backbones. With DeepSeek-v4-flash-0731, it achieves 93.69% symbolic accuracy on LSR-Transform, compared with 62.16% for SR-Scientist, and retains 72.97% accuracy on an anonymized variant that removes scientific descriptions and variable semantics, versus 39.64% for SR-Scientist. Under the same DeepSeek-v4-flash-0731 backbone, SRHarness also substantially outperforms Codex (72.97% vs. 20.72%) and reaches performance comparable to Codex with GPT-5.5, while simply providing Codex with the same scientific tools does not reproduce this advantage. These results show that effective agentic symbolic regression depends not only on models or tools, but also on structured runtime support for organizing scientific actions, accumulated hypotheses, and long-horizon search.