arXivPaperNeeds Review
SABRE: Scalable and Automated Benchmarking of VLMs under Stress
SABRE turns a Markdown test design into generated VLM stress tests, then filters and repairs candidates. It offers a repeatable pattern for refreshing evals as models improve.
arXiv
Source Summary
SABRE converts a Markdown task design into specifications, images, and question-answer pairs, then applies model filtering and human review. SABRE-Prior includes **600 images** and **1,000 questions** testing whether models follow visual evidence over learned expectations.
Practical Implication
Teams evaluating vision agents can encode test intent and schema first, generate candidates, discard easy cases, then reserve human effort for validity checks, annotation fixes, and localized image repair.
Agent-Ready Context
SABRE converts a Markdown task design into specifications, images, and question-answer pairs, then applies model filtering and human review. SABRE-Prior includes **600 images** and **1,000 questions** testing whether models follow visual evidence over learned expectations. Teams evaluating vision agents can encode test intent and schema first, generate candidates, discard easy cases, then reserve human effort for validity checks, annotation fixes, and localized image repair. Across **six VLMs**, macro-average accuracy ranged from **17.8% to 31.3%**. A real-image control was comparably difficult for the filtering model, so low scores cannot be attributed only to counterfactual generated imagery.
Context Map
benchmarkimage#agent-evals#benchmark-integrity#generative-mediaUncertainty
Across **six VLMs**, macro-average accuracy ranged from **17.8% to 31.3%**. A real-image control was comparably difficult for the filtering model, so low scores cannot be attributed only to counterfactual generated imagery.