Ledger:从第一视角视频构建持久化 3D 物体记忆
Never Look Back: Understanding Persistence in 3D Object Memory from Egocentric Videos
这篇论文提出 Ledger,让具身智能体能记住你放过的东西在哪,HD-EPIC 准确率从 29.7% 拉到 42.6%,做空间记忆方向可以细看。
Ledger 是一种持久化 3D 物体记忆方法,让具身助手通过观察人的第一视角视频建立物体记忆。它记录物体位置、历史与上下文描述,物体离开视野后仍保留记录,并通过静止位置聚类和多次移动证据来抑制定位噪声。在 HD-EPIC 基准上准确率从 29.7% 提升到 42.6%,UCS-Bench 从 33.8% 提升到 38.5%,在 Ego4D 上返回预测的中位定位误差为 0.99 米。研究还基于 100 条多场景拼接视频流分析了检索与构建两端的失败模式。
Never Look Back: Understanding Persistence in 3D Object Memory from Egocentric Videos
As we move through the world and carry out everyday tasks, we encounter objects that may become relevant only later. We are capable of recalling where we left something or what was inside a container, even without knowing we would need it later. Here, we study how an embodied assistant can build a similar memory from egocentric videos, by observing a person's day-to-day activities. We present Ledger, a persistent 3D object memory that combines object locations, their histories, and contextual descriptions. It associates observations across the recording and retains objects after they leave the view, including those the person never touches. It clusters each object's observations by resting locations and records a move only after repeated evidence, reducing the effect of localization noise. Short descriptions preserve details such as an object's contents or supporting surface. It saves these records to later answer spatial questions without having to access the original images or video. Our memory raises HD-EPIC accuracy from 29.7% to 42.6%, UCS-Bench accuracy from 33.8% to 38.5% and localizes Ego4D objects with a 0.99 m median error on returned predictions. Our analyses identify complementary roles for temporal persistence, contextual descriptions, and retrieval. Our study on 100 stitched streams of multiple scenes each further exposes failures in both retrieval and construction. Per-scene construction partially recovers the performance lost across scene changes compared to that of single scene streams.