论文

WorldSolver 基准测试:LLM 智能体能写物理求解器吗

WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?

精选理由

有人把 61 篇图形学论文的物理仿真做成了 168 道题考 LLM,GPT-5.6-Sol 也只考了 48.7 分,想看模型写仿真代码的真实水平可以读读。

研究团队发布 WorldSolver 基准,包含 168 个仿真任务,素材来自 61 篇经典计算机图形学论文,覆盖 7 个物理领域。每个任务给出固定的场景代码框架,由 LLM 智能体补全求解器实现。评估从 Execution Checks、Visual Fidelity 和 Physical Plausibility 三个维度进行。结果显示最强模型表现有限:GPT-5.6-Sol 总分 48.7%,Claude-Opus-5 为 46.7%,说明可执行求解器生成难度高,视觉与物理正确性更难满足。

原文 · arXiv cs.AI

WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?

LLM-based agents are increasingly advancing scientific and engineering problem solving, with physics simulation emerging as a challenging yet practical testbed for reproducing complex physical phenomena with application in embodied AI, games and films. As the workhorse of such simulation, a solver computes how the state of a dynamic system evolves over time. Building such solvers requires physical understanding to identify appropriate models, mathematical reasoning to formulate the underlying dynamics, and software engineering to implement them as executable code, yet this capability of LLM agents remains underexplored. To this end, we introduce WorldSolver, a benchmark of 168 simulation tasks derived from physical phenomena in 61 classic computer graphics papers, spanning 7 physical domains. Each task contains a code scaffold that provides a fixed simulation environment for the scene, with the solver implementation left for the agent to complete. Specifically, we evaluate them along three dimensions: Execution Checks for successful execution, Visual Fidelity for reproducing the intended dynamic behavior in the rendered simulation, and Physical Plausibility for physics-grounded verification of the generated dynamics. Experiments on frontier agents reveal that producing executable solvers is difficult itself, and satisfying visual and physical correctness is even harder. GPT-5.6-Sol and Claude-Opus-5 perform comparatively better than the other evaluated agents, yet achieve overall scores of only 48.7% and 46.7%, respectively. WorldSolver is an early step toward agentic solver generation, and we hope it helps drive progress toward agents that can faithfully simulate the dynamic physical world. Code is available at https://github.com/sirujiang/WorldSolver.