论文多源确认

Mine Odyssey 基准用 Minecraft 重建真实地点评测空间智能体

Mine Odyssey: Benchmarking Spatial Agentic Intelligence in the Wild

精选理由

拿 Minecraft 重建曼哈顿、白金汉宫这些真实地点来考模型找路,GPT-6 Astra 考了 85.6 分,开源的 DeepSeek-V4.1-Flash 才 23.9,差距很直观。

Mine Odyssey 是一个评测智能体空间智能的新基准,基于 Minecraft 重建的真实地点构建,包含 180 个任务、覆盖五大洲 20 个国家和地区的 30 个地点(20 个室外、10 个室内)。每个任务用自然语言指定要按顺序到达的路标点,智能体需寻找可行路线、开门、用楼梯和梯子跨层移动并从导航错误中恢复。在 8 个受测模型中,GPT-6 Astra 以 85.6% 成功率居首,Claude Opus 5.5 为 73.9%,开源权重模型中表现最好的 DeepSeek-V4.1-Flash 仅 23.9%。

原文 · arXiv: DeepSeek

Mine Odyssey: Benchmarking Spatial Agentic Intelligence in the Wild

Advances in foundation models are driving efforts to introduce agents to assist people in the physical world. Such agents require agentic spatial intelligence: exploring unfamiliar environments, updating spatial understanding through interaction, and adapting actions based on feedback to sustain progress toward a sequence of goals. Existing benchmarks cover only a limited range of spatial layouts, scales, and traversal requirements. We introduce Mine Odyssey, a benchmark for evaluating agentic spatial intelligence using Minecraft reconstructions of real-world locations. It comprises 180 tasks covering 30 such locations across 20 countries and regions on five continents, including 20 outdoor and 10 indoor settings. These settings span diverse spatial scales, layouts, terrains, and connectivity patterns, from Midtown Manhattan and rural Entrup to Santa Lucía Hill and Buckingham Palace. We select meaningful waypoints, such as landmarks, buildings, and rooms, and manually verify their accessibility. Each task provides a natural-language instruction specifying which waypoints to visit and in what order. Completing these tasks requires agents to find accessible routes and entrances, open doors, and move between levels using stairs and ladders, while monitoring their progress and recovering from navigation errors. Across eight evaluated state-of-the-art models, GPT-6 Astra achieves the highest success rate of 85.6%. However, the second-best model, Claude Opus 5.5, completes 73.9% of tasks, while the strongest evaluated open-weight model, DeepSeek-V4.1-Flash, reaches 23.9%, highlighting substantial room for improvement in the agentic spatial intelligence of current models. Comprehensive analyses and ablation studies on Mine Odyssey reveal current models' limitations and provide insights for advancing agentic spatial intelligence.