论文

Generative Tutorial:把生成式教程画进用户真实环境的AR系统

Generative Tutorial: Towards Live Contextualized Visual Instructions for Physical Tasks

精选理由

一篇arXiv论文:AR加图像生成,教程直接画在你自己的工作台上,24人实验显示比预制教程更管用。

论文提出 Generative Tutorial 框架,让图像与视频生成模型在用户自己的工作环境中实时呈现操作步骤和预期结果。团队先对当前 SOTA 图像与视频生成模型在 15 项实体任务上做形成性评估,找出失败模式与可用之处。据此构建的 AR 原型会读取工作台上下文、预测前一步动作的视觉结果,主动生成目标图像和演示视频。24 人实验室研究显示,相比预先制作的教程,该系统带来更高的任务完成质量、更强的环境对应感,以及更短的步骤确认间隔。定性分析还发现环境相似度影响用户信任,生成错误会干扰对指导的解读。

原文 · arXiv cs.AI

Generative Tutorial: Towards Live Contextualized Visual Instructions for Physical Tasks

Visual instructions for physical tasks are typically authored in one context and followed in another, requiring users to translate demonstrated tools, materials, and spatial relationships into their own environment. We introduce Generative Tutorial, a conceptual framework for live visual instruction that depicts intended outcomes and actions within the user's environment and task flow. A formative evaluation of state-of-the-art image and video generation identifies failures and potential benefits across 15 physical tasks. Drawing on these findings, we build an augmented-reality prototype system that proactively generates goal images and demonstration videos using observed workspace context and predicted visual outcomes of preceding actions. A 24-participant lab study found higher task performance quality, greater perceived workspace correspondence, and shorter step-confirmation intervals with the system than with pre-authored guidance. Qualitative findings highlighted how contextual resemblance shapes trust, how generation errors affect interpretation, and how guidance delivery should adapt to users' needs, informing future designs.