# The Alignment Illusion in Multimodal Large Language Models

Source: [arXiv](https://arxiv.org/abs/2609.30210v1)  
Feed7 permalink: https://feed7.dev/p/2609-30210v1-198in2y  
Published: 2026-09-24T17:42:29.000Z  
Trust: Needs Review (needs_review)

## Why Included

Common visual-text alignment scores stayed high after visual tokens were replaced with noise. Multimodal evaluations should pair internal geometry with controlled corruption and task accuracy.

## Source Summary

Across **13 multimodal models** from five families, replacing visual tokens with Gaussian noise sharply reduced accuracy. Yet **four standard alignment measures** did not consistently distinguish corrupted inputs from originals.

## Practical Implication

Builders evaluating vision systems should not read a scalar representation-similarity score as proof that image content is being integrated. Pair internal probes with controlled visual corruption and downstream task evidence.

## Agent-Ready Context

Across **13 multimodal models** from five families, replacing visual tokens with Gaussian noise sharply reduced accuracy. Yet **four standard alignment measures** did not consistently distinguish corrupted inputs from originals.

Builders evaluating vision systems should not read a scalar representation-similarity score as proof that image content is being integrated. Pair internal probes with controlled visual corruption and downstream task evidence.

The proposed **principal-angle gap** tracked accuracy more consistently under graded corruption, but structured irrelevant images still exposed cases where geometry and performance diverged. It remains a diagnostic, not a direct content-understanding score.

## Connected Context

Feed7 judgment across 875 accumulated Signals:

This narrows what internal alignment metrics establish for vision models: representational similarity may remain reassuring even after useful visual information is destroyed. Controlled corruption and task performance are therefore prerequisites for interpreting internal probes. The principal-angle gap improves diagnosis under graded noise but does not remove the need for behavioral evidence, especially with structured distractors.

- [DiaVLo: Diagnosing Behaviours of Vision-Language Models](https://feed7.dev/p/2609-22008v1-0w8qm00) — DiaVLo motivates internal causal diagnostics, while this result sets a boundary on such probes: internal geometry must be validated against controlled input interventions and downstream behavior.
- [SABRE: Scalable and Automated Benchmarking of VLMs under Stress](https://feed7.dev/p/2608-07435v1-0h6gzdk) — SABRE’s generated visual stress tests provide the kind of controlled behavioral evidence needed to supplement representation-level alignment measures.
- [The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping](https://feed7.dev/p/2608-06361v1-1n3dr85) — Both expose apparently favorable aggregate signals that can persist without faithful use of the input: alignment scores under corrupted images and final counts without accurate event recovery.
- [A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal](https://feed7.dev/p/2609-21996v1-16fmyv3) — PIR supports inspecting internal states when outputs are inconclusive; this study adds the complementary warning that an internal-state metric is not informative until interventions link it to task performance.

## Context Map

- Layer: benchmark
- Domains: image
- Topics: agent-evals, benchmark-integrity

## Uncertainty

- The proposed **principal-angle gap** tracked accuracy more consistently under graded corruption, but structured irrelevant images still exposed cases where geometry and performance diverged. It remains a diagnostic, not a direct content-understanding score.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
