Sign InOpen Brain
arXivPaperNeeds Review

DiaVLo: Diagnosing Behaviours of Vision-Language Models

DiaVLo supplements VLM scores with behavior specifications and causal concept estimates, helping builders see how a model reaches aligned or misaligned outputs.

arXiv · Sep 18, 2026
Open Source Open MarkdownOpen JSON
Source Summary

DiaVLo uses human curation and model generation to specify desired and observed VLM behavior, then surfaces mismatches. It also estimates which concepts most influence behavior and was evaluated on **several open-source VLMs** in **classification and generation** settings.

Practical Implication

For vision-model selection or evaluation, pair aggregate performance with behavioral labels and influential-concept analysis. This can expose how a model perceives, organizes, or prioritizes concepts when a score alone hides the failure mode.

Agent-Ready Context
DiaVLo uses human curation and model generation to specify desired and observed VLM behavior, then surfaces mismatches. It also estimates which concepts most influence behavior and was evaluated on **several open-source VLMs** in **classification and generation** settings.

For vision-model selection or evaluation, pair aggregate performance with behavioral labels and influential-concept analysis. This can expose how a model perceives, organizes, or prioritizes concepts when a score alone hides the failure mode.

The abstract says behavior labels correlate with performance but provides no effect sizes, model names, or benchmark counts. Human curation also means the usefulness of a diagnosis may depend on the quality of the behavioral specification.
Connected Context · Feed7 Judgment

DiaVLo adds a diagnostic layer that the prior benchmark-integrity work largely lacks: after controlling task and harness quality, evaluators can characterize how a VLM organizes concepts and why its behavior diverges from the specification. It complements rather than replaces leakage controls, realistic inputs, stress tests, and aggregate scores, and its reliance on curated behavior labels makes specification quality part of the evaluation.

SABRE: Scalable and Automated Benchmarking of VLMs under StressSABRE generates specification-led stress cases, while DiaVLo can characterize the behavioral and conceptual mismatch those cases expose; together they connect test construction with failure diagnosis.Can Edge-Deployable Vision-Language Models Identify Species?The species study shows deployment-image and open-set failures that aggregate clean-photo results miss; DiaVLo offers a way to label such behavioral differences and inspect the concepts influencing them.Stop Evaluating Models Like It's the 50s - Alejandro Vidal, MindmakersItem response theory diagnoses weak or informative questions at the suite level, whereas DiaVLo diagnoses model behavior and influential concepts, making the methods complementary rather than interchangeable.
Context Map
benchmarkimage#agent-evals#benchmark-integrity#model-selection
Uncertainty
The abstract says behavior labels correlate with performance but provides no effect sizes, model names, or benchmark counts. Human curation also means the usefulness of a diagnosis may depend on the quality of the behavioral specification.