HarnessSQL:训练与部署共用执行环境,Qwen3-14B 在 Spider 2.0 提升至 54.8%
阿里 Qwen3 模型在训练时套上部署用的执行环境,Spider 2.0 成绩直接翻倍多,做法是把数据库装进沙盒里边跑边学。
一篇论文提出 HarnessSQL 方法,核心是让模型在部署时使用的同一个执行 harness 内进行训练。Qwen3-14B 在 Spider 2.0-SQLite 上的成绩从 22.2% 提升到 54.8%,Qwen3-8B 从 15.5% 提升到 45.2%。传统 Text-to-SQL 模型只训练写单条静态查询,而部署后的数据库智能体需要检查 schema、运行探针查询并迭代修改。HarnessSQL 构建带隐藏答案校验的隔离可执行数据库,教师在目标 harness 内运行,只保留通过验证的轨迹做 SFT,再用执行奖励做 RL。两个模型的成绩还能迁移到 BIRD-Interact 和 LiveSQLBench 基准上。
Pay close attention to custom harnesses. This is a super interesting paper showing the impact of training inside the harness. Qwen3-14B goes from 22.2% to 54.8% on Spider 2.0-SQLite when it is trained in the same execution harness it uses at deployment. Text-to-SQL models are usually trained to write one static query. Deployed database agents inspect schemas, run probe queries and revise, and the harness for that only appears at inference time. HarnessSQL builds isolated executable databases with hidden answer checks. Teachers run inside the target harness; we keep only verified trajectories for SFT, then apply execution-reward RL. In addition, Qwen3-8B rises from 15.5% to 45.2%, and both models transfer to BIRD-Interact and LiveSQLBench. Paper: arxiv.org/abs/2610.12274 Chat with Paper: academy.dair.ai/papers/harness… 💬 4 🔄 1 ❤️ 14 👀 1228 📊 7 ⚡