Sign InOpen Brain
arXivPaperNeeds Review

The Alignment Illusion in Multimodal Large Language Models

Common visual-text alignment scores stayed high after visual tokens were replaced with noise. Multimodal evaluations should pair internal geometry with controlled corruption and task accuracy.

arXiv · Sep 24, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Across **13 multimodal models** from five families, replacing visual tokens with Gaussian noise sharply reduced accuracy. Yet **four standard alignment measures** did not consistently distinguish corrupted inputs from originals.

Practical Implication

Builders evaluating vision systems should not read a scalar representation-similarity score as proof that image content is being integrated. Pair internal probes with controlled visual corruption and downstream task evidence.

Agent-Ready Context
Across **13 multimodal models** from five families, replacing visual tokens with Gaussian noise sharply reduced accuracy. Yet **four standard alignment measures** did not consistently distinguish corrupted inputs from originals.

Builders evaluating vision systems should not read a scalar representation-similarity score as proof that image content is being integrated. Pair internal probes with controlled visual corruption and downstream task evidence.

The proposed **principal-angle gap** tracked accuracy more consistently under graded corruption, but structured irrelevant images still exposed cases where geometry and performance diverged. It remains a diagnostic, not a direct content-understanding score.
Connected Context · Feed7 Judgment

This narrows what internal alignment metrics establish for vision models: representational similarity may remain reassuring even after useful visual information is destroyed. Controlled corruption and task performance are therefore prerequisites for interpreting internal probes. The principal-angle gap improves diagnosis under graded noise but does not remove the need for behavioral evidence, especially with structured distractors.

DiaVLo: Diagnosing Behaviours of Vision-Language ModelsDiaVLo motivates internal causal diagnostics, while this result sets a boundary on such probes: internal geometry must be validated against controlled input interventions and downstream behavior.SABRE: Scalable and Automated Benchmarking of VLMs under StressSABRE’s generated visual stress tests provide the kind of controlled behavioral evidence needed to supplement representation-level alignment measures.The Low Frequency Trap: Video Language Models Fail at Simple Event BookkeepingBoth expose apparently favorable aggregate signals that can persist without faithful use of the input: alignment scores under corrupted images and final counts without accurate event recovery.A Lie Detector Test for Language Models: Reading Knowledge a Model Won't RevealPIR supports inspecting internal states when outputs are inconclusive; this study adds the complementary warning that an internal-state metric is not informative until interventions link it to task performance.
Context Map
benchmarkimage#agent-evals#benchmark-integrity
Uncertainty
The proposed **principal-angle gap** tracked accuracy more consistently under graded corruption, but structured irrelevant images still exposed cases where geometry and performance diverged. It remains a diagnostic, not a direct content-understanding score.