BrickBench:评估智能体乐高搭建能力的基准发布
BrickBench: Evaluating Agentic Brick Design
有人做了个让 AI 拼乐高的评测基准,智能体得选零件还得保证能真的拼出来,结果还是拼不过人类,挺有意思。
BrickBench 是一个面向文本条件乐高套装设计的智能体评测基准,要求智能体从离散零件库中选件并生成可实际拼装的模型。基准从 validity、alignment 和 design 三个维度打分,覆盖三种不同规模和零件可用性的设置。论文同时发布 BrickAgent 环境,供编程智能体构建、检查和验证设计。结果显示主流智能体能满足可验证的物理与语义要求,但与人类设计仍有差距。
BrickBench: Evaluating Agentic Brick Design
We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about local and global constraints. We score validity, alignment, and design across three settings that vary in scale and part availability. We provide BrickAgent, an environment for coding agents to construct, inspect, and validate their designs. We find that leading agents largely satisfy verifiable physical and semantic requirements, but fall short of human designs. We release our benchmark and environment at http://www.brickben.ch