SOTA Generative Media Panel — Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind
DeepMind’s panel shows why generative-media evals need task-specific human review: broad preferences can miss repeated artifacts, exact sizing, text errors, and brand consistency.
The panel says **Nano Banana 2 Light** targets faster, cheaper generation and editing, with roughly **3-second latency**. It also describes a human test where regenerated scenes were broadly preferred to real-video counterparts.
Builders should evaluate media models on their actual production constraints: exact text, repeated patterns, object scale, reference consistency, audio-visual behavior, and brand colors. Side-by-side review remains necessary when models are close.
The panel says **Nano Banana 2 Light** targets faster, cheaper generation and editing, with roughly **3-second latency**. It also describes a human test where regenerated scenes were broadly preferred to real-video counterparts. Builders should evaluate media models on their actual production constraints: exact text, repeated patterns, object scale, reference consistency, audio-visual behavior, and brand colors. Side-by-side review remains necessary when models are close. Broad preference scores can conceal unusable details and learned artifacts, such as recurring wedding rings on hands. The speakers also leave the right intermediate representation—language, code, continuous tokens, or something else—as an **open question**.
The panel reinforces that fast, inexpensive media generation does not make broad preference scores sufficient for production selection. Its artifact examples and unresolved representation question narrow evaluation toward constraint-specific, side-by-side inspection of text, patterns, scale, references, color, and audio-visual behavior. The reported human preference result therefore signals promise without resolving whether outputs meet exact production requirements.