模型多源确认

Anthropic Haiku 5.5 跑分 Browser Use Bench v2.1:性能与 Luna 持平,价格高出 2.8 倍

精选理由

Browser Use 拿 Haiku 5.5 跑了自家基准 v2.1,性能和 Luna 打平但贵 2.8 倍,长任务下还触发 5 倍加价,选型前可以看看这份实测。

Browser Use 团队公布了 Anthropic Haiku 5.5 在 Browser Use Bench v2.1 上的成绩,性能与 Luna 持平,但价格贵 2.8 倍。该基准以超长时程任务为主,Haiku 5.5 在 100k context 后价格升至 5 倍,缓存 token 也计入计费。测试中 Haiku 5.5 频繁触发这一高价档位。

原文 · Browser Use

Haiku 5.5 on Browser Use Bench v2.1 👀 > same performace as Luna > 2.8x more expensive Haiku gets 5x more expensive after 100k context, including cached tokens. BU Bench 2.1 is mostly extremely long-running horizon tasks. It hits that pricing tier a lot. 💬 5 🔄 1 ❤️ 30 👀 2186 📊 11 ⚡