Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs
Vision encoders can expose an object’s typical color even from grayscale input, and VLM post-training can substantially alter that signal. Useful evidence that visual representations contain learned concepts, not only pixels.
Researchers probed vision encoders with color and grayscale objects. An object’s **canonical color remained decodable from grayscale images**, and that signal was connected to predicted object identity.
For builders evaluating visual agents, pixel-level tests may miss conceptual associations already present in the encoder. Canonical-color probes offer a controlled way to compare what object semantics remain linearly accessible before and after VLM post-training.
Researchers probed vision encoders with color and grayscale objects. An object’s **canonical color remained decodable from grayscale images**, and that signal was connected to predicted object identity. For builders evaluating visual agents, pixel-level tests may miss conceptual associations already present in the encoder. Canonical-color probes offer a controlled way to compare what object semantics remain linearly accessible before and after VLM post-training. The study uses one constrained concept as its lens and reports no quantitative results in the supplied material. It shows decodability, not that a VLM will reliably use the concept in an application. The paper was accepted to **EMNLP 2026**.
This adds a representation-level diagnostic to visual-agent evaluation: a model can retain an object-associated concept even when the corresponding pixels are absent. It therefore separates what an encoder makes linearly accessible from what a downstream VLM actually uses, narrowing any behavioral failure claim that attributes the problem simply to missing visual information.