arXivPaperNeeds Review
MIRROR: Learning from the Other View for Multi-Modal Reasoning
MIRROR trains a VLM across text, diagram, and combined views by letting its strongest view supervise weaker ones, targeting the modality inconsistency that single-view evals hide.
arXiv · Jul 23, 2026
Source Summary
ODA-Data pairs geometry problems across **text-dominant**, **image-dominant**, and **combined image+text** views. MIRROR evaluates every view, selects the best-performing one as teacher, then aligns the others toward it with a reverse-KL objective.
Practical Implication
Builders of multimodal agents should test semantically equivalent inputs in each supported modality. Divergent answers expose failures that aggregate accuracy hides, while reciprocal supervision offers a way to reuse the model's strongest representation.
Agent-Ready Context
ODA-Data pairs geometry problems across **text-dominant**, **image-dominant**, and **combined image+text** views. MIRROR evaluates every view, selects the best-performing one as teacher, then aligns the others toward it with a reverse-KL objective. Builders of multimodal agents should test semantically equivalent inputs in each supported modality. Divergent answers expose failures that aggregate accuracy hides, while reciprocal supervision offers a way to reuse the model's strongest representation. The reported gains cover geometry reasoning benchmarks and comparisons with standard RL, but the supplied material includes no scores. It remains unclear whether the method transfers to noisier screenshots, documents, or mixed-modal coding tasks.
Context Map
modelimageresearch#reasoningUncertainty
The reported gains cover geometry reasoning benchmarks and comparisons with standard RL, but the supplied material includes no scores. It remains unclear whether the method transfers to noisier screenshots, documents, or mixed-modal coding tasks.