LiveMACE:评估LLM智能体在动态市场中的能力
LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets
LiveMACE用真实市场数据评估LLM智能体能力,发现结果与能力不匹配,比单纯看结果更准确。
LiveMACEBench基准测试使用实时金融市场作为持续LLM智能体的自然演化测试环境。研究对五个前沿LLM模型进行了30天的持续评估,测试了工具使用、持久记忆、规则遵循和多智能体协作配置。研究发现结果与能力之间存在显著差距:实际回报往往与能力特定测量值不同,相似结果可能源于机制使用的不同模式。
LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets
Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions. We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents. Five frontier LLMs operate along continuous trajectories under matched Tool Use, Persistent Memory, Rule Following, and Multi-Agent Collaboration configurations. We evaluate them through both realized outcomes and mechanism-specific diagnostics derived from complete decision traces. Across 30 days of live evaluation, we find a pronounced outcome-capability gap: realized returns often diverge from capability-specific measurements, and similar outcomes can arise from markedly different patterns of mechanism use. Trace-level diagnostics further expose distinct bottlenecks across capabilities, demonstrating that mechanism access, effective mechanism use, and downstream performance are not interchangeable measures of agent capability. LiveMACEBench makes this distinction measurable, turning live markets from a performance leaderboard into a diagnostic environment for agent capability