# Gemini 3.8 Flash now available on AI Gateway

Source: [Vercel](https://vercel.com/changelog/gemini-3-8-flash-now-available-on-ai-gateway)  
Feed7 permalink: https://feed7.dev/p/gemini-3-8-flash-now-available-on-ai-gateway-0cun77q  
Published: 2026-09-02T00:00:00.000Z  
Trust: Official Source (official_source)

## Why Included

Gemini 3.8 Flash brings multimodal input, tool calling, web search, and default reasoning to coding agents through Vercel. Its temporary 50% discount runs through December 31.

## Source Summary

Google’s **Gemini 3.8 Flash** is now on Vercel AI Gateway with a **1M-token context window** and a 65,536-token output limit. It accepts text, images, PDFs, and video, and supports tool calls and web search.

## Practical Implication

Builders can select google/gemini-3.8-flash in connected coding agents. Thinking is enabled by default, so test its latency, token use, and behavior on representative repository tasks before changing a production routing default.

## Agent-Ready Context

Google’s **Gemini 3.8 Flash** is now on Vercel AI Gateway with a **1M-token context window** and a 65,536-token output limit. It accepts text, images, PDFs, and video, and supports tool calls and web search.

Builders can select google/gemini-3.8-flash in connected coding agents. Thinking is enabled by default, so test its latency, token use, and behavior on representative repository tasks before changing a production routing default.

Vercel says it improves software engineering, agent work, and multi-step reasoning at the prior Flash model’s speed and cost, but supplies no evaluation results here. The advertised **50% discount ends December 31**.

## Connected Context

Feed7 judgment across 669 accumulated Signals:

This adds another million-token, multimodal Flash route to the same coding-agent model pool, but does not resolve selection among neighboring models. Its tool calls, web search, and large output allowance broaden the test surface; the absent evaluation data and default thinking make repository-matched measurements of quality, latency, and token use the deciding evidence, with the temporary discount separated from durable routing economics.

- [Qwen 3.8 Flash now available on AI Gateway](https://feed7.dev/p/qwen-3-8-flash-now-available-on-ai-gateway-1skoa7y) — Qwen 3.8 Flash has the same stated context and output limits, making it a close matched candidate for testing multimodal coding, tool use, latency, price, and reliability rather than choosing by capacity.
- [GLM 5.3 now available on AI Gateway](https://feed7.dev/p/glm-5-3-now-available-on-ai-gateway-0s7o9zv) — GLM 5.3 reinforces that a 1M-token window and provider claims do not distinguish repository-scale routes without controlled workload evaluation.
- [The Base Model Is Dead — Varun Singh, Arcee AI](https://feed7.dev/p/the-base-model-is-dead-varun-singh-arcee-ai-02hts76) — The discussion of capability differences established during model training explains why similar gateway specifications may still yield materially different agent behavior.

## Context Map

- Layer: model
- Domains: coding
- Topics: coding-agents, reasoning, model-selection

## Uncertainty

- Vercel says it improves software engineering, agent work, and multi-step reasoning at the prior Flash model’s speed and cost, but supplies no evaluation results here. The advertised **50% discount ends December 31**.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
