论文

ViGeo:用时空类比把视觉上下文学习扩展到视频任务

Video, Ergo Genero: Unifying Video Tasks via Spatiotemporal Analogy

精选理由

训练免了!ViGeo 用时空类比让视频模型靠示例学新任务,还能零样本处理 event cameras,顺手拆了个提示词捷径的坑。

研究团队提出 ViGeo 框架,通过时空画布补全(spatiotemporal canvas completion)把视觉类比式上下文学习从图像扩展到视频。该方法免训练即可适配多种视频任务,并能零样本泛化到未见过的视频操作和 event cameras 等新模态。论文还发现任务内化现象:预训练任务的查询格式会覆盖示例,加入少量任务无关数据即可消除这一捷径。

原文 · arXiv cs.AI

Video, Ergo Genero: Unifying Video Tasks via Spatiotemporal Analogy

Adapting video models to new tasks typically requires dedicated data curation and fine-tuning. While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain. To explore whether analogy-based methods can unify diverse video tasks and generalize to out-of-distribution scenarios, we introduce ViGeo, a framework that extends visual in-context learning to the video domain via spatiotemporal canvas completion. Evaluated on a diverse task taxonomy with a strict train-test split, ViGeo generalizes to unseen video manipulations and zero-shot modalities (e.g., event cameras). Finally, we identify task internalization, where a query format associated with a pretrained task overrides the demonstration, and show that this shortcut can be removed with a small amount of task-unrelated data, highlighting the need to decorrelate prompt format from task identity.