EnigmaForge 基准:不给问题只给故事,25 个模型直觉能力差 22 倍
EnigmaForge: The Question Is Hidden in the Story
25 个模型在这个基准上只给故事不给问题,直觉成绩差了 22 倍,排名和常规榜单完全对不上,还有模型被自己的过滤器拦住,很值得细看。
EnigmaForge 是一个把逻辑谜题藏进文档叙事里的基准,由 SAT 求解器在生成时验证唯一解,并附带消融证书确保每条线索都不可或缺。25 个前沿模型在 600 多个实例(17,400 条评分记录)上测试,主指标是只看故事时的直觉成功率,世界重建是次级指标。结果显示直觉排名与事实恢复排名大幅错位:事实恢复只有 1.6 倍差距,直觉却有 22 倍差距,事实恢复第二名在直觉榜上只排第十四。部分模型还被自家内容过滤器拦下,无法到达谜题本身,说明把拒绝计为失败的基准实际测的是过滤器行为。
EnigmaForge: The Question Is Hidden in the Story
Most benchmarks hand the model a question. EnigmaForge hands it a stack of old documents and no question at all. Buried in the letters, receipts, and logbook margins is a small logic puzzle whose solution is unique - proved by a SAT solver at generation time, with an ablation certificate showing every clue is load-bearing. Because instances are generated rather than collected, the corpus renews forever. The headline measure is intuition: task success when handed only the story, with world reconstruction as the secondary axis. Twenty-five frontier models ran over 600 instances (17,400 scored records) under three matched conditions. Intuition reshuffles the leaderboard: a 22x spread where fact recovery spans 1.6x, the second-best fact-recoverer ranks fourteenth, one model is indifferent to being told the question, and another is significantly better without it. Several models were blocked by their own content filters before reaching the puzzle - any benchmark scoring refusals as failure is quietly measuring filter behavior.