论文

Strategy-Diverse RL 训练开源 LLM 做工业级优化元求解器

LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization

精选理由

用强化学习把开源模型训练成会挑求解策略的优化器,成绩还能压过 GPT-5.5,做运筹优化的可以看看。

论文提出 Strategy-Diverse Reinforcement Learning(SDRL),把开源 LLM 训练成自适应的优化元求解器。实验显示求解器集成推理、精确组合算法和启发式搜索在不同问题结构与规模上各有互补优势。SDRL 通过正确性门控的分层多样性奖励,防止策略过早坍缩。训练方案混合支持纯文本问题和基于文件的真实实例。在基准测试和工业级优化任务上,该框架超过了 DeepSeek-V4-Pro 和 GPT-5.5 等前沿模型。

原文 · arXiv: DeepSeek

LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization

Scaling LLM-based optimization from textbook-scale instances to real-world, industrial tasks remains a critical open challenge. Existing approaches are predominantly evaluated on small, self-contained textual problems and often commit to a solver-integrated paradigm, limiting their ability to handle the scale and structural diversity of practical optimization workloads. In this work, we propose a practical framework for training open-source LLMs to tackle real-world, industrial-scale optimization. We first show empirically that solver-integrated reasoning, exact combinatorial algorithm, and heuristic search exhibit complementary strengths across different problem structures and scales. Motivated by this, we introduce Strategy-Diverse Reinforcement Learning (SDRL), which trains LLMs as adaptive optimization meta-solvers. SDRL leverages this complementarity through a correctness-gated hierarchical diversity reward that promotes robust exploration across varying strategies and within each strategy, effectively preventing premature strategy collapse. We further introduce a mixed-format training scheme that jointly supports both self-contained textual problems and file-grounded instances. Across comprehensive evaluations, our framework outperforms existing fine-tuned methods and frontier models including DeepSeek-V4-Pro and GPT-5.5, both on average across benchmarks and on industrial-scale optimization tasks.