# SOTA Generative Media Panel — Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind

Source: [AI Engineer](https://www.youtube.com/watch?v=KLDdXOw6jIc)  
Feed7 permalink: https://feed7.dev/p/sota-generative-media-panel-dumitru-erhan-shane-gu-nicole-brichtova-goog-1yskl7e  
Published: 2026-08-30T14:00:06.000Z  
Trust: Source Linked (source_linked)

## Why Included

DeepMind’s panel shows why generative-media evals need task-specific human review: broad preferences can miss repeated artifacts, exact sizing, text errors, and brand consistency.

## Source Summary

The panel says **Nano Banana 2 Light** targets faster, cheaper generation and editing, with roughly **3-second latency**. It also describes a human test where regenerated scenes were broadly preferred to real-video counterparts.

## Practical Implication

Builders should evaluate media models on their actual production constraints: exact text, repeated patterns, object scale, reference consistency, audio-visual behavior, and brand colors. Side-by-side review remains necessary when models are close.

## Agent-Ready Context

The panel says **Nano Banana 2 Light** targets faster, cheaper generation and editing, with roughly **3-second latency**. It also describes a human test where regenerated scenes were broadly preferred to real-video counterparts.

Builders should evaluate media models on their actual production constraints: exact text, repeated patterns, object scale, reference consistency, audio-visual behavior, and brand colors. Side-by-side review remains necessary when models are close.

Broad preference scores can conceal unusable details and learned artifacts, such as recurring wedding rings on hands. The speakers also leave the right intermediate representation—language, code, continuous tokens, or something else—as an **open question**.

## Connected Context

Feed7 judgment across 617 accumulated Signals:

The panel reinforces that fast, inexpensive media generation does not make broad preference scores sufficient for production selection. Its artifact examples and unresolved representation question narrow evaluation toward constraint-specific, side-by-side inspection of text, patterns, scale, references, color, and audio-visual behavior. The reported human preference result therefore signals promise without resolving whether outputs meet exact production requirements.

- [Evaling Video Slop — Maor Bril, Character.ai](https://feed7.dev/p/evaling-video-slop-maor-bril-character-ai-0cd76sd) — Both show that polished or preferred output can conceal decisive failures; the video-evaluation candidate adds the need for time-aware checks of action, physics, and storytelling.
- [Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber](https://feed7.dev/p/building-closed-loop-evals-for-a-multimodal-agent-at-scale-soumya-gupta-1cqjbe2) — Uber’s golden sets, iterative QA, and production feedback provide an implementation pattern for the panel’s call to evaluate exact constraints and catch recurring artifacts.
- [SABRE: Scalable and Automated Benchmarking of VLMs under Stress](https://feed7.dev/p/2608-07435v1-0h6gzdk) — SABRE offers a repeatable way to generate and refresh targeted visual stress tests, complementing the panel’s recommendation to probe specific production failure modes rather than rely on aggregate preference.
- [Start building with Nano Banana 2 Lite and Gemini Omni Flash](https://feed7.dev/p/gemini-omni-flash-nano-banana-2-lite-0uezrjl) — The API announcement supplies concrete latency and price context for the cheaper-generation direction, while the panel explains why those operating metrics still need constraint-level quality review.

## Context Map

- Layer: benchmark
- Domains: image, video
- Topics: generative-media, benchmark-integrity

## Uncertainty

- Broad preference scores can conceal unusable details and learned artifacts, such as recurring wedding rings on hands. The speakers also leave the right intermediate representation—language, code, continuous tokens, or something else—as an **open question**.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
