Artificial Analysis 编码代理索引新增安全拒绝报告
Artificial Analysis 新增了安全拒绝报告,能看模型什么时候拒绝、切换到哪个备用模型,Claude Sonnet 拒绝率只有 Opus 的一半。
Artificial Analysis 编码代理索引新增安全拒绝报告功能,可显示拒绝发生时间和使用的备用模型。Claude Code with Sonnet 5.5 (max) 目前排名第一,安全拒绝率为 4.5%,约为 Claude Code with Opus 5.5 (max) 8.9% 的一半。Sonnet 5.5 的 94% 拒绝发生在首次交互后,拒绝后几乎总是切换到 Opus 4.8。
Safety refusal reporting in the Artificial Analysis Coding Agent Index now shows when refusals occur and which models are used as fallbacks
We've added two ways to explore the results:
➤ Refusal timing: see whether a refusal occurred from the task prompt alone or later in the task, after the agent had already started working
➤ Fallback model: see which model the agent switched to after a refusal, alongside attempts that were blocked
Claude Code with Sonnet 5.5 (max) currently ranks first in the Index. Its safety refusal rate is 4.5%, roughly half the 8.9% recorded for Claude Code with Opus 5.5 (max).
Around 94% of Sonnet 5.5's refusals occurred after the first turn. Following a refusal, the agent almost always switched to Opus 4.8.
Observed fallback patterns vary across model and agent configurations. With Fable 5.1, Claude Code fell back predominantly to Opus 4.8, while Opus 5 accounted for a much larger share of Devin Fusion's fallbacks in the Index.
Both views are available for the overall Index and each benchmark. Safety and fallback behavior is provider-configured and may change over time; these rates reflect behavior recorded at the time of benchmarking.