AI代理基准测试与实际应用案例
The scenes we keep replaying: - Laurie Voss reran a year old benchmark that used to break models at...
看看AI代理在基准测试和实际应用中的表现,了解当前AI代理的优缺点。
Laurie Voss重新运行了一年前的基准测试,过去模型在200条指令时会失败,现在所有模型都能立即获得100%分数,上限已接近2000。Cornelia Davis指出目前没有代理支持MCP任务,因为规范在不断变化。Dustin Mihalik在Indeed的搜索中添加了小部件,Claude停止深入挖掘,文本结果获得了15次搜索和精选表格,而小部件只获得了一次调用。
The scenes we keep replaying: - Laurie Voss reran a year old benchmark that used to break models at...
The scenes we keep replaying: - Laurie Voss reran a year old benchmark that used to break models at 200 instructions. Every current model scored 100 percent instantly. The ceiling now sits near 2,000. Total research bill: $29. - Cornelia Davis says no agent supports MCP tasks yet because the maintainers are smart. The spec keeps changing. Alex Hancock opened the next talk. He maintains one, and it's just because he's lazy. - Dustin Mihalik gave Indeed's job search a widget and Claude stopped digging. Text results got fifteen searches and a cherry picked table. The widget got one call. Nobody wants ten carousels. - Google's demo agent hit a database error, decided the fix was to delete the table and start fresh, and did it. Nothing stopped it. - Atlan changed its positioning. The sales agent on Atlan's own website kept pitching the old version. No one had told it. 💬 0 🔄 0 ❤️ 0 👀 49 ⚡