视觉编码器与视觉语言模型概念解码研究
Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs
这篇论文揭示了视觉编码器如何存储物体概念信息,即使在没有颜色的情况下也能识别物体的标准颜色。
研究人员通过标准颜色作为测试案例,探究视觉编码器在灰度图像中保留颜色概念信息的能力。研究构建了一个包含标准颜色物体的数据集,使用彩色和灰度图像探测视觉编码器对颜色和物体身份的解码能力。研究发现,即使输入图像中没有颜色,标准颜色信息仍可从灰度图像中解码,并与预测的物体身份相关联,表明存在概念性联系。
Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs
Visual encoders construct a representation of the image input for Vision-Language models. How much conceptual, as opposed to immediately visible, information does this representation contain? We use canonical color as a controlled test case to ask whether vision encoders make canonical-color information linearly accessible, even when color is removed from the input image. We construct a dataset of objects with canonical colors, and probe vision encoders for both color and object identity using color and grayscale images. We find that canonical color remains decodable from grayscale images, and is tied to predicted object identity, indicating a conceptual link. Extending this analysis to full VLMs, we find that VLM post-training can have a surprisingly large effect on color decodability in the vision encoder. Overall, canonical color provides a usefully controllable lens for tracing object-level conceptual semantic information in vision encoders and VLMs.