The Alignment Illusion in Multimodal Large Language Models
Common visual-text alignment scores stayed high after visual tokens were replaced with noise. Multimodal evaluations should pair internal geometry with controlled corruption and task accuracy.
Across **13 multimodal models** from five families, replacing visual tokens with Gaussian noise sharply reduced accuracy. Yet **four standard alignment measures** did not consistently distinguish corrupted inputs from originals.
Builders evaluating vision systems should not read a scalar representation-similarity score as proof that image content is being integrated. Pair internal probes with controlled visual corruption and downstream task evidence.
Across **13 multimodal models** from five families, replacing visual tokens with Gaussian noise sharply reduced accuracy. Yet **four standard alignment measures** did not consistently distinguish corrupted inputs from originals. Builders evaluating vision systems should not read a scalar representation-similarity score as proof that image content is being integrated. Pair internal probes with controlled visual corruption and downstream task evidence. The proposed **principal-angle gap** tracked accuracy more consistently under graded corruption, but structured irrelevant images still exposed cases where geometry and performance diverged. It remains a diagnostic, not a direct content-understanding score.
This narrows what internal alignment metrics establish for vision models: representational similarity may remain reassuring even after useful visual information is destroyed. Controlled corruption and task performance are therefore prerequisites for interpreting internal probes. The principal-angle gap improves diagnosis under graded noise but does not remove the need for behavioral evidence, especially with structured distractors.