Sign InOpen Brain
arXivPaperNeeds Review

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

Controlled traces show video models can improve final counting scores without faithfully recovering events, so agents handling video need timestamp-level checks, not answer-only evals.

arXiv
Open Source Open MarkdownOpen JSON
Source Summary

The study profiles event counting across **2,190 controlled videos** with executable traces. Gemini 3.6 Flash reached 80% reliability for persistent transitions up to **12 events**, but had no reliable positive-count region for transient blinks.

Practical Implication

For video agents, evaluate the reported event sequence against timestamps, not just the final count. In the high-count, high-frequency regime, only **0.2%** of counts were correct and **18.1%** of true events were recovered.

Agent-Ready Context
The study profiles event counting across **2,190 controlled videos** with executable traces. Gemini 3.6 Flash reached 80% reliability for persistent transitions up to **12 events**, but had no reliable positive-count region for transient blinks.

For video agents, evaluate the reported event sequence against timestamps, not just the final count. In the high-count, high-frequency regime, only **0.2%** of counts were correct and **18.1%** of true events were recovered.

More frames raised Bounce Ball accuracy from 19.6% to 29.3%, while full sequence agreement remained 3.7%. Prompt changes also brought limited gains, so higher aggregate accuracy may conceal poor event bookkeeping.
Context Map
benchmarkvideo#agent-evals#benchmark-integrity#agent-reliability
Uncertainty
More frames raised Bounce Ball accuracy from 19.6% to 29.3%, while full sequence agreement remained 3.7%. Prompt changes also brought limited gains, so higher aggregate accuracy may conceal poor event bookkeeping.