BOTTLED 基准:LLM 智能体能否把能力封装成低成本的专用方案
Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
这篇论文测试了让智能体自己动手省钱:让 Opus 5 等模型把任务封装成小程序或小模型,成本能降 657 倍但多数尝试会翻车,结果挺有意思。
论文提出 bottling 概念,指 LLM 智能体把通用能力转化为任务专用的低成本方案,例如训练小模型或写可复用程序。BOTTLED 基准给智能体完整无标注工作负载,要求在固定时间、算力和 API 预算内完成。跨 10 个模型、3 个任务的测试显示,60 次运行中有 48 次低于各自模型零样本表现的 95% 置信区间下界,31 次跑不过同等 token 预算下的小模型蒸馏基线。但收益也很可观:Opus 5 在查询-商品相关性分类上以约 657 倍更低的成本保留约 82% 的零样本 macro-F1,并以四分之一成本恢复专用廉价模型 Jev 约 94% 的 macro-F1。
Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
Large language models (LLMs) can solve many narrow tasks, but querying them separately for millions of related instances can be prohibitively expensive. Can LLM agents autonomously create cheaper solutions for such workloads? We call this ability "bottling": the ability to turn general capabilities into task-specific solutions that balance answer quality and amortised cost. We introduce BOTTLED, a benchmark in which agents receive an entire unlabelled workload and must complete it under fixed time, compute and LLM API budgets. Agents choose their own approach, such as training a small model or writing a reusable program. Across ten models and three tasks, we find that strong zero-shot task performance does not reliably translate into strong bottling capabilities. Models with similar zero-shot scores can differ substantially after bottling, and 48 of 60 bottling runs score below the lower bound of the 95% confidence interval of their model's zero-shot performance. Moreover, 31 of 60 runs underperform the stronger of two small-model distillation baselines with the same token budget. Nevertheless, bottling can yield substantial savings: on query-product relevance classification, Opus 5 retains about 82% of its zero-shot macro-F1 at roughly 657 times lower reported cost. Bottling is also competitive with Jev, a "system one" model built especially for cheap, repetitive inference: Opus 5 on the same task recovers about 94% of Jev's macro-F1 at a quarter of Jev's projected full-workload cost. BOTTLED provides a basis for evaluating and improving agents' ability to invest limited resources in reusable solutions for large, repetitive workloads.
- rohanpaul_ai10-06 21:16原文