{
  "schema_version": "1.1",
  "id": "s13:https://arxiv.org/abs/2609.30210v1",
  "slug": "2609-30210v1-198in2y",
  "url": "https://feed7.dev/p/2609-30210v1-198in2y",
  "title": "The Alignment Illusion in Multimodal Large Language Models",
  "why_included": "Common visual-text alignment scores stayed high after visual tokens were replaced with noise. Multimodal evaluations should pair internal geometry with controlled corruption and task accuracy.",
  "summary": "Across **13 multimodal models** from five families, replacing visual tokens with Gaussian noise sharply reduced accuracy. Yet **four standard alignment measures** did not consistently distinguish corrupted inputs from originals.",
  "practical_implication": "Builders evaluating vision systems should not read a scalar representation-similarity score as proof that image content is being integrated. Pair internal probes with controlled visual corruption and downstream task evidence.",
  "agent_context": "Across **13 multimodal models** from five families, replacing visual tokens with Gaussian noise sharply reduced accuracy. Yet **four standard alignment measures** did not consistently distinguish corrupted inputs from originals.\n\nBuilders evaluating vision systems should not read a scalar representation-similarity score as proof that image content is being integrated. Pair internal probes with controlled visual corruption and downstream task evidence.\n\nThe proposed **principal-angle gap** tracked accuracy more consistently under graded corruption, but structured irrelevant images still exposed cases where geometry and performance diverged. It remains a diagnostic, not a direct content-understanding score.",
  "source": {
    "name": "arXiv",
    "url": "https://arxiv.org/abs/2609.30210v1",
    "published_at": "2026-09-24T17:42:29.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Paper",
  "layer": "benchmark",
  "domains": [
    "image"
  ],
  "topics": [
    "agent-evals",
    "benchmark-integrity"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "The proposed **principal-angle gap** tracked accuracy more consistently under graded corruption, but structured irrelevant images still exposed cases where geometry and performance diverged. It remains a diagnostic, not a direct content-understanding score."
  ],
  "connected_context": {
    "meaning": "This narrows what internal alignment metrics establish for vision models: representational similarity may remain reassuring even after useful visual information is destroyed. Controlled corruption and task performance are therefore prerequisites for interpreting internal probes. The principal-angle gap improves diagnosis under graded noise but does not remove the need for behavioral evidence, especially with structured distractors.",
    "corpus_size": 875,
    "generated_at": "2026-09-25T09:07:42.999Z",
    "connections": [
      {
        "title": "DiaVLo: Diagnosing Behaviours of Vision-Language Models",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2609.22008v1",
        "feed7_url": "https://feed7.dev/p/2609-22008v1-0w8qm00",
        "reason": "DiaVLo motivates internal causal diagnostics, while this result sets a boundary on such probes: internal geometry must be validated against controlled input interventions and downstream behavior."
      },
      {
        "title": "SABRE: Scalable and Automated Benchmarking of VLMs under Stress",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2608.07435v1",
        "feed7_url": "https://feed7.dev/p/2608-07435v1-0h6gzdk",
        "reason": "SABRE’s generated visual stress tests provide the kind of controlled behavioral evidence needed to supplement representation-level alignment measures."
      },
      {
        "title": "The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2608.06361v1",
        "feed7_url": "https://feed7.dev/p/2608-06361v1-1n3dr85",
        "reason": "Both expose apparently favorable aggregate signals that can persist without faithful use of the input: alignment scores under corrupted images and final counts without accurate event recovery."
      },
      {
        "title": "A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2609.21996v1",
        "feed7_url": "https://feed7.dev/p/2609-21996v1-16fmyv3",
        "reason": "PIR supports inspecting internal states when outputs are inconclusive; this study adds the complementary warning that an internal-state metric is not informative until interventions link it to task performance."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-09-24T17:42:29.000Z",
  "modified_at": "2026-09-24T17:42:29.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/2609-30210v1-198in2y",
    "json": "https://feed7.dev/p/2609-30210v1-198in2y.json",
    "markdown": "https://feed7.dev/p/2609-30210v1-198in2y.md"
  }
}