# Gemini 3.8 Live models now available on AI Gateway

Source: [Vercel](https://vercel.com/changelog/gemini-3-8-live-models-now-available-on-ai-gateway)  
Feed7 permalink: https://feed7.dev/p/gemini-3-8-live-models-now-available-on-ai-gateway-14yndjj  
Published: 2026-09-15T00:00:00.000Z  
Trust: Official Source (official_source)

## Why Included

Two Gemini 3.8 Live models add real-time audio to AI Gateway; the extended variant can reason alongside speech while the base model keeps tool calls in the background.

## Source Summary

AI Gateway now exposes **Gemini 3.8 Live** and **Gemini 3.8 Live Extended Thinking** through the AI SDK's realtime API. Both handle spoken interaction; the base model adds visual grounding, background tool calls, and switching across **97 languages**.

## Practical Implication

For voice agents, use short-lived tokens and the supplied WebSocket adapter, then choose the extended model only when multi-step reasoning must continue alongside speech. Gateway can centralize usage, cost, retries, and failover.

## Agent-Ready Context

AI Gateway now exposes **Gemini 3.8 Live** and **Gemini 3.8 Live Extended Thinking** through the AI SDK's realtime API. Both handle spoken interaction; the base model adds visual grounding, background tool calls, and switching across **97 languages**.

For voice agents, use short-lived tokens and the supplied WebSocket adapter, then choose the extended model only when multi-step reasoning must continue alongside speech. Gateway can centralize usage, cost, retries, and failover.

The material gives no latency, pricing, or quality measurements, and realtime support is exposed through an **experimental API**. Extended Thinking also requires choosing exactly one thinking control: level or budget.

## Connected Context

Feed7 judgment across 807 accumulated Signals:

This broadens realtime voice routing with visual grounding, background tools, multilingual switching, and an explicit choice between ordinary and extended reasoning. Against existing live models, it makes concurrent speech-and-work capability a selectable architecture rather than a unique feature. Missing latency, cost, and quality measurements—and the experimental API—leave model choice dependent on matched voice workloads.

- [Grok Voice Think Fast 2.0 now available on AI Gateway](https://feed7.dev/p/grok-voice-think-fast-2-0-now-available-on-ai-gateway-1dacr27) — Grok is a direct realtime speech-and-tool competitor; together they make interruption behavior, tool-call timing, noisy-input handling, credential design, and measured latency relevant comparison points.
- [GPT-Live 1 now available on AI Gateway](https://feed7.dev/p/gpt-live-1-now-available-on-ai-gateway-18x7ban) — Both support speech alongside deeper work, but GPT-Live-1 delegates that work to a separately chosen text model while Gemini offers an Extended Thinking live variant, creating distinct architectures to evaluate.
- [Gemini 3.8 Flash now available on AI Gateway](https://feed7.dev/p/gemini-3-8-flash-now-available-on-ai-gateway-0cun77q) — Gemini 3.8 Flash already makes thinking behavior an evaluation variable for latency and token use; the live release sharpens that requirement by forcing exactly one extended-thinking control while speech continues.

## Context Map

- Layer: model
- Domains: audio
- Topics: reasoning, tool-use

## Uncertainty

- The material gives no latency, pricing, or quality measurements, and realtime support is exposed through an **experimental API**. Extended Thinking also requires choosing exactly one thinking control: level or budget.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
