Sign InOpen Brain
AI EngineerVideoSource Linked

Voice agents with Realtime Video — Sidney Primas, LemonSlice

LemonSlice’s avatar stack treats long-running visual stability, audio-conditioned emotion, and deterministic action timing as the core engineering problems beyond lip sync.

AI Engineer · Aug 18, 2026
Open Source Open MarkdownOpen JSON
Source Summary

LemonSlice’s deployed Roosevelt avatar generates continuously for **8 hours without a reset**, while another deployment targets 16 hours. Its API accepts a single image and sits above customers’ own LLM and voice stacks; pricing is described as similar to a voice model.

Practical Implication

For video agents, plan around error accumulation and behavioral control, not only frame quality. Audio data and encoders affect expression, while an emotion engine must time gestures and reactions against incoming audio and response text.

Agent-Ready Context
LemonSlice’s deployed Roosevelt avatar generates continuously for **8 hours without a reset**, while another deployment targets 16 hours. Its API accepts a single image and sits above customers’ own LLM and voice stacks; pricing is described as similar to a voice model.

For video agents, plan around error accumulation and behavioral control, not only frame quality. Audio data and encoders affect expression, while an emotion engine must time gestures and reactions against incoming audio and response text.

The company does not claim to have passed its proposed avatar Turing test. Its more natural model was not real time in the demo, and the method used to suppress long-horizon errors was not disclosed.
Context Map
toolsvideoaudio#generative-media#agent-reliability
Uncertainty
The company does not claim to have passed its proposed avatar Turing test. Its more natural model was not real time in the demo, and the method used to suppress long-horizon errors was not disclosed.