Sign InOpen Brain
OpenAIOfficial ReleaseOfficial Source

Build more natural voice experiences with GPT‑Live‑1 in the API

GPT-Live-1 adds full-duplex voice, improved instruction following, custom voices, and telephony support for builders shipping conversational audio products.

OpenAI · Sep 10, 2026
Open Source Open MarkdownOpen JSON
Source Summary

**GPT-Live-1** brings **full-duplex voice conversations** to the API, alongside stronger instruction following, custom voices, and telephony support.

Practical Implication

Voice-agent builders should reconsider architectures that separate turn detection, response generation, and playback when a live, bidirectional API can handle the conversational loop more directly.

Agent-Ready Context
**GPT-Live-1** brings **full-duplex voice conversations** to the API, alongside stronger instruction following, custom voices, and telephony support.

Voice-agent builders should reconsider architectures that separate turn detection, response generation, and playback when a live, bidirectional API can handle the conversational loop more directly.

The material provides no latency, pricing, language coverage, safety controls, or migration details, so production tradeoffs cannot yet be assessed from it.
Connected Context · Feed7 Judgment

This shifts the voice-agent design choice from assembling transcription, turn handling, generation, and playback toward a managed full-duplex conversational loop. It confirms that bidirectional speech APIs are becoming a distinct model surface, but missing latency, interruption, safety, language, and cost evidence prevents deciding whether the integrated loop should replace modular audio pipelines.

Grok Voice Think Fast 2.0 now available on AI GatewayBoth expose realtime bidirectional voice, making interruption timing, latency, tool-call behavior, and client credential handling relevant comparison dimensions rather than text-model quality alone.Gemini 3.5 Transcribe now available on AI GatewayIts live transcription route represents the modular input stage that GPT-Live-1 may absorb, while retaining explicit multilingual and custom-vocabulary capabilities not specified for the integrated API.Fish Audio models now available on Vercel AI Gateway for freeFish Audio supports separately composed transcription and speech-generation pipelines, providing an architectural contrast to GPT-Live-1’s single full-duplex conversational loop.Voice agents with Realtime Video — Sidney Primas, LemonSliceThe avatar-stack case shows that an integrated voice loop would still leave visual stability, audio-conditioned emotion, and deterministic action timing as separate engineering concerns for embodied agents.
Context Map
modelaudio#generative-media#agent-sdks
Uncertainty
The material provides no latency, pricing, language coverage, safety controls, or migration details, so production tradeoffs cannot yet be assessed from it.