# Evaling Video Slop — Maor Bril, Character.ai

Source: [AI Engineer](https://www.youtube.com/watch?v=b_PmGocP4rc)  
Feed7 permalink: https://feed7.dev/p/evaling-video-slop-maor-bril-character-ai-0cd76sd  
Published: 2026-07-25T00:00:02.000Z  
Trust: Source Linked (source_linked)

## Why Included

Video evaluators can reward polish while missing frozen action, broken physics, or failed storytelling. Builders need time-aware criteria and human-calibrated data, not frame quality alone.

## Source Summary

A first evaluator gave **9.2 for camera work** to a clip whose camera stayed still for **4 seconds**. Frame-level metrics caught appearance and drift, but missed whether motion, physics, pacing, sound, and story matched the intended video.

## Practical Implication

Train and calibrate evaluators around the exact quality axis you need. **Pairwise comparisons** proved easier to align than numeric ratings, and checking short generations before assembly can prevent wasted rendering and editing.

## Agent-Ready Context

A first evaluator gave **9.2 for camera work** to a clip whose camera stayed still for **4 seconds**. Frame-level metrics caught appearance and drift, but missed whether motion, physics, pacing, sound, and story matched the intended video.

Train and calibrate evaluators around the exact quality axis you need. **Pairwise comparisons** proved easier to align than numeric ratings, and checking short generations before assembly can prevent wasted rendering and editing.

Manufactured defects taught the first model to recognize gloss and artificial artifacts instead of quality. Mixing real and generated footage can instead produce an AI detector, while lip-sync evaluation remains unresolved and human-calibrated judging is slow and costly.

## Context Map

- Layer: benchmark
- Domains: video
- Topics: agent-evals, benchmark-integrity, generative-media

## Uncertainty

- Manufactured defects taught the first model to recognize gloss and artificial artifacts instead of quality. Mixing real and generated footage can instead produce an AI detector, while lip-sync evaluation remains unresolved and human-calibrated judging is slow and costly.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
