Sign InOpen Brain
AI EngineerVideoSource Linked

Evaling Video Slop — Maor Bril, Character.ai

Video evaluators can reward polish while missing frozen action, broken physics, or failed storytelling. Builders need time-aware criteria and human-calibrated data, not frame quality alone.

AI Engineer · Jul 25, 2026
Open Source Open MarkdownOpen JSON
Source Summary

A first evaluator gave **9.2 for camera work** to a clip whose camera stayed still for **4 seconds**. Frame-level metrics caught appearance and drift, but missed whether motion, physics, pacing, sound, and story matched the intended video.

Practical Implication

Train and calibrate evaluators around the exact quality axis you need. **Pairwise comparisons** proved easier to align than numeric ratings, and checking short generations before assembly can prevent wasted rendering and editing.

Agent-Ready Context
A first evaluator gave **9.2 for camera work** to a clip whose camera stayed still for **4 seconds**. Frame-level metrics caught appearance and drift, but missed whether motion, physics, pacing, sound, and story matched the intended video.

Train and calibrate evaluators around the exact quality axis you need. **Pairwise comparisons** proved easier to align than numeric ratings, and checking short generations before assembly can prevent wasted rendering and editing.

Manufactured defects taught the first model to recognize gloss and artificial artifacts instead of quality. Mixing real and generated footage can instead produce an AI detector, while lip-sync evaluation remains unresolved and human-calibrated judging is slow and costly.
Context Map
benchmarkvideo#agent-evals#benchmark-integrity#generative-media
Uncertainty
Manufactured defects taught the first model to recognize gloss and artificial artifacts instead of quality. Mixing real and generated footage can instead produce an AI detector, while lip-sync evaluation remains unresolved and human-calibrated judging is slow and costly.