论文

IBIB协议:评估企业AI系统的新方法

IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier

精选理由

IBIB协议解决了企业AI系统评估的测量误差问题,提供了新的评估方法和算法,对AI系统开发者很有参考价值。

IBIB协议通过服务路由而非模型标识符来测量企业AI系统。该协议包含三个部分:黄金盲能力绑定预检、可靠性包含的第一轮评分规则和结构盲评分裁决。研究团队在11个系统上测试了该协议,发现能力可用性是可测量的,且系统间的区分度不均匀。服务臂选择将一个声明的修订版本和精度从77.38提升到82.54,配对区间为[0.11,10.60]。

原文 · arXiv cs.LG

IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier

Enterprises deploy systems, not checkpoints. Usable capability depends jointly on weights, serving route, precision, output contract, and harness, yet all 18 audited benchmarks score advertised model identifiers. We treat this as measurement error and give a protocol that makes it reportable. It has three parts. A gold-blind capability-binding preflight verifies that a route can execute the evaluation contract before any task reaches it; a reliability-inclusive first-pass scoring rule keeps failure in the score while keeping unsupported capability out; and adjudication is structurally score-blind. We call the protocol IB2 and release its algorithms, classification tables, request contract, and manifest schemas. Its reference instantiation, 128 locked tasks and 987 assertions over document, spreadsheet, chart, tool and database work, stays sealed: the procedure is the artifact, not the corpus. Across eleven systems, four results. Capability availability is measurable: two complete single-route runs on identical weights later failed distinct predicates of the finalized binding gate, while a third passed that gate before a fresh run. The advertised identifier exposed neither limit. Discrimination is not uniform: four of seven suites saturate under a six-system band, with the spread almost entirely from governed database work and multi-tab joins, so we report interval-backed resolution groups, not ranks; two of the nominal five-label output's four cuts fail multiplicity adjustment. Serving-arm choice moved one declared revision and precision from 77.38 to 82.54, paired interval [0.11,10.60], though the arms differ in access mode, harness generation, and the serving tool-call parser, and harness generation is a property of our evaluator, not any endpoint. Excluding failed responses from denominators changes the point ordering, so reliability inclusion changes a conclusion, not its wording.