Gary Marcus 批评 OpenAI 数学证明结果报告缺乏关键细节
Gary Marcus 炮轰 OpenAI 的数学证明报告:没写过程、没写失败率,啥关键信息都没有,这篇吐槽把问题列得很清楚。
Gary Marcus 在 X 上批评 OpenAI 发布的数学证明结果报告。他指出报告没有说明证明是一次生成后由 Lean 验证,还是经过迭代过程,也没披露失败率、训练与数据增强方法。Marcus 认为这种详细程度通不过同行评审,无法判断该结果能否泛化到数学之外,可能是通向 AGI 的一步,也可能只是 Lean 和合成数据在可验证领域里的巧妙应用。
The real news here isn’t the result; it’s what we were not told. 1. AI once tried to be a science. Now we get stuff like the completely vague report from OpenAI below, and a lot of ignorant questions from people who don’t know how to think critically. “Same procedure”? “using an unreleased model”? This would never pass peer review. We don’t know what the procedure was. We know nothing about the architecture (e.g., were proofs generated in one shot, and then verified by Lean? was there an iterative process?). We know nothing about the failure rate. We know nothing about the training/post training/data agumentation. 2. As a result we have zero idea of how generalizable the result is outside math. 3. A lot of X has been reduced to an ignorant cheering section that applauds without knowing what it is applauding or what it might mean — without ever asking basic scientific questions. The new system could be a legitimate step towards AGI or just a clever leveraging of Lean and synthetic data in a verifiable domain with no generality whatsoever; from this report we can tell almost nothing. 💬 22 🔄 10 ❤️ 92 👀 5311 📊 31 ⚡