腾讯推出新模型 EvolveScaler,通过逻辑状态机生成可执行文本
🚀 EvolveScaler is here. Read a 40-day RPG log. Now answer one question: if you skip the mini-boss ...
腾讯新出的 EvolveScaler 模型,能通过逻辑状态机生成可执行文本,比普通模型在 8 个基准测试上平均提升 5.25 分。
腾讯推出的 EvolveScaler 模型通过定义世界为可执行状态机,然后将其渲染为自然语言。该模型在 8 个外分布基准测试上平均提升 5.25 分,在 14 个前沿模型中最难级别上,中位数平均 @5 降至 11.3。
🚀 EvolveScaler is here. Read a 40-day RPG log. Now answer one question: if you skip the mini-boss ...
🚀 EvolveScaler is here. Read a 40-day RPG log. Now answer one question: if you skip the mini-boss on Day 7, do you still beat the final boss? The answer isn't in the log. You have to replay the world. That's Information Evolution — records get retracted, corrected, backfilled. The world keeps changing after you read it. So we build it backwards: define the world as an executable state machine, then render it into natural language. Code guarantees the logic. Language delivers the mess. ➡️ 117 prototypes. 159 question operators. 5 difficulty tiers. Up to ~1,200 events per sample. ➡️ 14 frontier models, hardest tier: median avg @5 falls to 11.3. ➡️ Train on it instead: +5.25 average across 8 out-of-distribution benchmarks. Check out our paper and project page. 📚 Paper: arxiv.org/abs/2609.08435 re 🏠 Project Pag tencent-hunyuan.github.io/evolve-scaler/ tCp 💬 1 🔄 1 ❤️ 28 👀 1834 📊 7 ⚡