论文

Endless Exam 基准:用14类数学构造衡量模型通往超级智能的进度

The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence

精选理由

有人出了个数学基准 Endless Exam,14类构造题自动验证打分,8个模型全测了一遍,目前没谁超过人类已发表纪录。

arXiv 论文提出 Endless Exam 基准,用14个参数化数学构造任务族衡量当前模型向超级智能的进展。每个提交对象自动验证有效性,并对照已发表前沿或构造基线给出相对质量分,分数不封顶于1。在69个不同实例上评估8个模型,连续质量分数能区分表现差异,但没有一个被测系统超过已发表的前沿结果。任务族取材于开放数学问题,可在更大参数下生成新实例,紧凑证书让大型构造仍可机器验证。作者已发布生成器、验证器、参考构造、模型响应与分析数据。

原文 · arXiv cs.AI

The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence

We introduce the Endless Exam, a benchmark for measuring mathematical progress from today's models toward artificial superintelligence through fourteen parameterised construction families. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at $1$. The families draw on open mathematical problems for long-term targets and generate new instances at larger parameters, where compact certificates keep large constructions verifiable. Across eight models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though no evaluated system surpasses a published frontier. Size-quality curves show how construction quality changes as problem size increases. We release the generators, verifiers, references, model responses and analysis to support continued measurement before and beyond human frontiers.