论文精选

量子工程验证自主性研究:新框架与基准测试揭示当前代理性能差异

Evaluating Verified Autonomy in Quantum Engineering

精选理由

研究团队用新框架和49个任务测试了17个量子代理,发现它们在可靠操作上差异很大,这个基准测试很有参考价值。

为评估量子工程中代理的验证自主性,研究者开发了Quantum-Harbor虚拟实验室,用于直接验证动作与结论。基于此框架,他们引入了QIQCBench基准测试,包含49个专家任务,涵盖校准、纠错等多个层面。在17个前沿代理系统上测试后,结果显示性能差异显著,暴露出当前能力与可靠操作之间的巨大差距。

原文 · arXiv cs.AI

Evaluating Verified Autonomy in Quantum Engineering

Reliable quantum engineering is essential for turning quantum phenomena into practical technologies. As quantum platforms grow in scale and complexity, their characterization and operation require increasing human effort and coordination. Scientific artificial intelligence agents, which can plan experiments, operate instruments, and analyze observations, offer a promising route towards autonomous quantum engineering. Yet whether current agents can perform reliably in this setting has not been systematically established. To fill this gap, we developed Quantum-Harbor, a virtual laboratory that provides a controlled execution environment for agents to interact with quantum systems. This design enables direct verification of both the actions taken and the conclusions drawn. Building on this framework, we introduce QIQCBench, a benchmark of $49$ expert-authored tasks spanning multiple layers including calibration and control, error correction and compilation, sensing and networking. Across $17$ frontier agentic systems, QIQCBench reveals wide variation in verified performance. These results expose a substantial gap between demonstrating capability and achieving reliable operation, and establish Quantum-Harbor as a foundation for measuring progress towards verified autonomy in quantum engineering.