论文

VideoMSN:图像分类器高效自监督视频表示学习

Image Classifiers are Efficient Self-Supervised Video Representation Learners

精选理由

VideoMSN用图像分类器做视频学习,比传统方法快32倍,还能在少样本分类中表现出色。

VideoMSN是一种掩码孪生网络框架,通过将视频表示为超图像(由视频帧组成的网格)实现高效自监督时空表示学习。该方法使用共享的Vision Transformer编码器,通过空间块掩码和时间帧掩码两种视图对齐嵌入,无需重建即可捕捉运动和外观线索。在Kinetics-400、UCF101和HMDB51基准上达到最先进性能,相比前人方法视频预训练epoch减少32倍和160倍。

原文 · arXiv cs.LG

Image Classifiers are Efficient Self-Supervised Video Representation Learners

We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learning with unlabeled data, we repurpose standard image Vision Transformers by representing videos as super images which are grids composed of frames sampled from videos. From each super image, we construct two views: one with spatial patch masking and the other with temporal frame masking, ensuring no information leakage across frames. A shared Vision Transformer (ViT) encoder aligns their embeddings using a masked Siamese loss, capturing both motion and appearance cues without reconstruction. Our decoder-free formulation leverages an image foundation model towards efficient video representation learning. Starting from pretrained DINO-v3 and DeiT-v3 image encoders, VideoMSN achieves state-of-the-art performance on Kinetics-400, UCF101, and HMDB51 while requiring up to $32\times$ fewer and $160\times$ fewer video pretraining epochs compared to prior video self-supervised learning methods. Our proposed approach also shows strong performance in low-shot classification, confirming the transferability of the learned representations in a label-scarce scenario. Project Page: https://cvir.github.io/projects/videomsn.