{
  "schema_version": "1.1",
  "id": "archive:https://arxiv.org/abs/2609.09124v1",
  "slug": "2609-09124v1-1bzksj0",
  "url": "https://feed7.dev/p/2609-09124v1-1bzksj0",
  "title": "Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs",
  "why_included": "Vision encoders can expose an object’s typical color even from grayscale input, and VLM post-training can substantially alter that signal. Useful evidence that visual representations contain learned concepts, not only pixels.",
  "summary": "Researchers probed vision encoders with color and grayscale objects. An object’s **canonical color remained decodable from grayscale images**, and that signal was connected to predicted object identity.",
  "practical_implication": "For builders evaluating visual agents, pixel-level tests may miss conceptual associations already present in the encoder. Canonical-color probes offer a controlled way to compare what object semantics remain linearly accessible before and after VLM post-training.",
  "agent_context": "Researchers probed vision encoders with color and grayscale objects. An object’s **canonical color remained decodable from grayscale images**, and that signal was connected to predicted object identity.\n\nFor builders evaluating visual agents, pixel-level tests may miss conceptual associations already present in the encoder. Canonical-color probes offer a controlled way to compare what object semantics remain linearly accessible before and after VLM post-training.\n\nThe study uses one constrained concept as its lens and reports no quantitative results in the supplied material. It shows decodability, not that a VLM will reliably use the concept in an application. The paper was accepted to **EMNLP 2026**.",
  "source": {
    "name": "arXiv",
    "url": "https://arxiv.org/abs/2609.09124v1",
    "published_at": "2026-09-08T17:50:09.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Paper",
  "layer": "benchmark",
  "domains": [
    "image"
  ],
  "topics": [
    "agent-evals"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "The study uses one constrained concept as its lens and reports no quantitative results in the supplied material. It shows decodability, not that a VLM will reliably use the concept in an application. The paper was accepted to **EMNLP 2026**."
  ],
  "connected_context": {
    "meaning": "This adds a representation-level diagnostic to visual-agent evaluation: a model can retain an object-associated concept even when the corresponding pixels are absent. It therefore separates what an encoder makes linearly accessible from what a downstream VLM actually uses, narrowing any behavioral failure claim that attributes the problem simply to missing visual information.",
    "corpus_size": 713,
    "generated_at": "2026-09-09T10:12:53.401Z",
    "connections": [
      {
        "title": "Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2607.09654v1",
        "feed7_url": "https://feed7.dev/p/2607-09654v1-0b5dedg",
        "reason": "The decade-spanning study measures behavioral visual-cognitive errors, while canonical-color probes offer a controlled internal signal that may help localize whether object semantics remain encoded."
      },
      {
        "title": "Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2609.01604v1",
        "feed7_url": "https://feed7.dev/p/2609-01604v1-02vljon",
        "reason": "Both move beyond aggregate scores by probing intermediate representations, separating accessible evidence from the later mechanism that converts it into a judgment."
      },
      {
        "title": "SABRE: Scalable and Automated Benchmarking of VLMs under Stress",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2608.07435v1",
        "feed7_url": "https://feed7.dev/p/2608-07435v1-0h6gzdk",
        "reason": "SABRE supplies scalable behavioral stress tests, whereas canonical-color probing can test whether a failure reflects absent encoder information or failure to use an available concept."
      },
      {
        "title": "Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2608.16868v1",
        "feed7_url": "https://feed7.dev/p/2608-16868v1-00830as",
        "reason": "Both caution that decodable internal information does not establish natural downstream use: accessibility of a signal is weaker evidence than its causal role in an output."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-09-08T17:50:09.000Z",
  "modified_at": "2026-09-08T17:50:09.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/2609-09124v1-1bzksj0",
    "json": "https://feed7.dev/p/2609-09124v1-1bzksj0.json",
    "markdown": "https://feed7.dev/p/2609-09124v1-1bzksj0.md"
  }
}