40亿参数本地模型作为候选生成器:机电调试中的外部验收层验证协议
Requirement-Bound Verified Commissioning: A Frozen Four-Billion-Parameter Local Model as a Candidate Generator under an External Acceptance Layer with Verification and Release Authority
机电调试场景里怎么防大模型编造计划?这篇论文的做法挺有意思:小模型只管提候选,发布权交给外部验证层,21 个伪造计划全被拦下,但也暴露了防不住错误用户答案的短板。
arXiv 论文提出将候选生成与发布权限分离的机电调试验收协议,需发布计划的模型仅有 40 亿参数且本地冻结运行。在 144 个任务的基准中,22 个无法回答的任务上模型伪造了 21 个就绪计划,全部被外部验证层拒绝。83 次发布中未观察到错误发布,Clopper-Pearson 单侧 95% 上界为 0.0354,低于封存的 5% 阈值。但基准外的 146 次发布中出现 1 次错误发布,且 431 组配对中有 169 组发布了错误用户答案,质疑策略与真实用户行为尚未测试。
Requirement-Bound Verified Commissioning: A Frozen Four-Billion-Parameter Local Model as a Candidate Generator under an External Acceptance Layer with Verification and Release Authority
An acceptance protocol is developed for sensor-coordinate and polarity binding in mechatronic commissioning. Candidate generation is separated from release authority. Requirements unsupported by a deterministic parser are routed to a frozen local language model with four billion parameters. Plans are released only when both facts can be derived by an external gate under a sealed grammar. One canonical answer is requested from a gold-standard user when eligible. The protocol was evaluated once under a criterion fixed before benchmark construction, on 144 tasks written by isolated agent contexts without access to the gate, grammar, or experimental plan. Three contributions are established. First, candidate generation and release decisions were measured separately. Fabricated ready plans were committed on 21 of 22 routed unanswerable tasks, and all were rejected. The same 83 releases were reproduced without model calls. Second, no false release was observed among 83 releases. A one-sided 95% Clopper-Pearson upper bound of 0.0354 was obtained as a diagnostic under an independent-and-identically-distributed assumption, below the sealed 5% threshold. However, one false release was subsequently recorded among 146 releases outside the benchmark at seed 0. Third, protection against incorrect user answers was characterized. Both facts were bound from the original text on 13 of 96 answerable tasks. Incorrect answers were released in 169 of 431 pairings on the remaining tasks, including failures involving coordinate exclusion. A deployable questioning policy was not tested because eligibility was determined from the answer key. Gate sensitivity and real user behavior were not measured.