Sign InOpen Brain
Atlas / Infra

Gateways

Open JSONConfidence: Auto-collectedLast updated 2026-09-02

Current Answer

No editorial synthesis yet — the evidence below is collected automatically from source labels. A current answer lands here once an editor approves one.

Evidence

GLM-5.3 is 50% off through DigitalOcean on AI Gateway
Vercel · 2026-09-02

GLM-5.3 is half-price through September 8 via a temporary DigitalOcean-only model ID. Keep the standard ID in durable agent configs if you need fallback after the offer.

x402 isn’t good (yet) — Jan Curn, Apify
AI Engineer · 2026-09-01

x402 servers can perform work before payment settlement, leaving a double-spend window. Builders should settle first or accept explicit counterparty risk until stronger schemes mature.

When AI Agents Pay and Sellers Monetize: Building x402 Apps on AWS — Anil Nadiminti, AWS
AI Engineer · 2026-09-01

AWS is separating agent payment policy from model execution: AgentCore handles wallets and limits, while WAF meters bot access. The useful pattern is deterministic spend control at the edge.

Why Your AI Agent Needs a Wallet: USDC and Nanopayments — Harshal Bhangale, Circle
AI Engineer · 2026-09-01

A wallet-equipped agent crossed paywalls and completed paid email and phone actions under a spending cap. The engineering lesson is to enforce budgets in the wallet, not in prompts.

Qwen 3.8 Max 0902 now available on AI Gateway
Vercel · 2026-09-01

Vercel’s gateway now exposes a pinned Qwen snapshot aimed at larger coding projects and longer agent runs, with a dated model ID that prevents silent upgrades.

Claude Fable 5.1 now available on AI Gateway
Vercel · 2026-09-01

Claude Fable 5.1 reaches Vercel AI Gateway with ordered fallbacks for classifier refusals, but its 30-day retention policy rules out zero-data-retention workloads.

Set per-user budgets on AI Gateway
Vercel · 2026-08-31

AI Gateway can now cap each user's aggregate spend across attributed API keys and app tokens, giving unattended coding agents a hard cost boundary.

MiniMax H3 and H3 Max are 50% off on AI Gateway
Vercel · 2026-08-30

Vercel is halving AI Gateway charges for MiniMax H3 and H3 Max through September 13; existing model IDs receive the discount without code changes.

Muse Image now available on AI Gateway
Vercel · 2026-08-26

Muse Image is available through Vercel AI Gateway for both generation and instruction-based editing, including reference-image guidance through the AI SDK.

Gemini 3.5 Transcribe now available on AI Gateway
Vercel · 2026-08-26

Gemini 3.5 Transcribe adds batch and live WebSocket transcription to AI Gateway, with automatic language detection, 85+ languages, and custom vocabulary.

The end of credential sprawl for agents
Vercel · 2026-08-25

Vercel Connect gives agents runtime-minted, task-scoped credentials instead of stored provider tokens, adding per-user identity, revocation, audit logs, and usage visibility.

MiniMax M3 and M2.7 are free on AI Gateway
Vercel · 2026-08-25

AI Gateway offers temporary free routes for MiniMax M3 and M2.7, but the -free model IDs hard-fail after September 6; standard IDs preserve provider fallback at normal rates.

Preferences Over Benchmarks: Model Routing — Archana Kamath & Tyler Gillam, DigitalOcean
AI Engineer · 2026-08-22

Per-task model routing cut the demonstrated coding session’s cost from 44¢ to 14¢ with similar completion time, but builders still need workload-specific evals to validate quality.

GPT-5.6 Sol is now 50% off a lower price
Vercel · 2026-08-21

GPT-5.6 Sol now costs less across Vercel AI Gateway tiers, with an additional 50% discount through September 18 and no model-ID change for existing agent setups.

Fish Audio models now available on Vercel AI Gateway for free
Vercel · 2026-08-19

Vercel AI Gateway added four Fish Audio models for speech generation and transcription, with AI SDK 7 support and a free window whose model naming determines later billing.

Security Firewall for Agents — Ryan Dahl, Deno
AI Engineer · 2026-08-17

Deno treats production agents as untrusted and filters their outbound traffic outside the agent, showing how broad operational access can coexist with protocol-aware controls.

GPT-5.6 Sol is 50% off on AI Gateway for the next month
Vercel · 2026-08-17

Vercel cut GPT-5.6 Sol pricing by 50% through September 18, making direct AI Gateway runs cheaper across every tier without changing model IDs or agent configs.

Gemini 3.7 Flash now available on AI Gateway for 50% off
Vercel · 2026-08-13

Gemini 3.7 Flash is on Vercel AI Gateway at 50% off through 2026, with direct setup paths for major coding agents and controls for routing, retries, and spend.

GLM 5.2 free for eve agents through August 27 via Blackbox on AI Gateway
Vercel · 2026-08-13

Eve agents can use GLM 5.2 free through August 27 via Vercel AI Gateway; new agents default to it, while existing agents need a model-setting change.

Set up coding agents in one command with AI Gateway
Vercel · 2026-08-12

Vercel’s setup command can route nine coding-agent clients through one gateway for shared models, budgets, policy, and traces. Review the in-place config edits before adopting it.

Grok Imagine Image 2.0 now available on Vercel AI Gateway
Vercel · 2026-08-08

Vercel AI Gateway now exposes xAI’s image model through the AI SDK, including 1K/2K generation, batches, and targeted edits that aim to preserve untouched details.

Vercel AI Gateway and Vercel Sandbox now available on Hermes Agent
Vercel · 2026-08-07

Hermes can route inference through Vercel’s model gateway and move command execution into an opt-in cloud microVM, separating model access from the machine where the agent runs.

Seedance 2.5 now available on Vercel AI Gateway
Vercel · 2026-08-06

Seedance 2.5 adds multimodal video generation and local edits to AI Gateway, with clips up to 30 seconds and separate image and video references.

Export AI Gateway traces with Vercel Drains
Vercel · 2026-08-05

Vercel AI Gateway can export per-request OpenTelemetry traces, exposing routing, retries, latency, tokens, cost, and attribution without sending prompt or completion content.

AI Gateway is now available on AWS Marketplace
Vercel · 2026-08-05

Vercel AI Gateway is available through AWS Marketplace, letting teams place inference on their AWS bill under annual private offers without changing per-token pricing.

DeepSeek V4 Flash is 90% off through Novita on AI Gateway
Vercel · 2026-08-04

Vercel Pro users can route DeepSeek V4 Flash to Novita for 90% off through August 11. Provider fallback remains enabled, but fallback requests use standard rates.

Qwen 3.8 Max now available on Vercel AI Gateway
Vercel · 2026-08-02

Vercel AI Gateway now exposes Qwen 3.8 Max to coding agents, adding one model endpoint for long-context text and vision work with gateway routing, budgets, and usage tracking.

AI Gateway now supports team and project spend budgets
Vercel · 2026-07-31

AI Gateway can now enforce spend caps across a team, project, or API key, giving agent workloads layered cost controls instead of relying on per-key limits alone.

AI Gateway logs now have a dedicated page
Vercel · 2026-07-31

AI Gateway’s dedicated logs expose per-request cost, tokens, latency, routing, and provider fallbacks, making agent failures and spend anomalies easier to trace.

10x more capacity for Laguna S 2.1 on AI Gateway
Vercel · 2026-07-31

AI Gateway has raised Laguna S 2.1 capacity tenfold for both paid and free model IDs, reducing throughput constraints for high-volume or long-running coding agents.

AI Gateway: GPT-5.6 pricing and speed updates
Vercel · 2026-07-30

Vercel cut Luna and Terra token prices and raised Sol fast-mode speed without changing model IDs, so existing agent workloads inherit the changes without code edits.

MiniMax H3 now available on AI Gateway
Vercel · 2026-07-30

MiniMax H3 brings short 2K video generation to Vercel AI Gateway, with text, keyframe, and multimodal reference inputs. Reference and keyframe modes cannot be combined.

AI Gateway adds unified fast mode support
Vercel · 2026-07-29

AI Gateway now exposes one beta fast-mode option across models, letting coding agents request lower latency while retaining standard-speed fallback when no fast tier exists.

Regional inference now available on AI Gateway
Vercel · 2026-07-27

AI Gateway can pin inference and retained provider data to the US or EU, giving agent builders one residency control with per-response verification across supported providers.

WebSocket support for OpenAI Responses API live on AI Gateway
Vercel · 2026-07-27

Vercel’s gateway now supports persistent WebSocket sessions for the Responses API, reducing repeated context transfer during long, tool-heavy agent runs.

Kimi K3 and Kimi K3 Fast with ZDR and US-based providers now on AI Gateway
Vercel · 2026-07-27

Kimi K3 now has US-hosted, ZDR-capable gateway routes plus a faster tier, giving coding-agent users explicit latency, residency, retention, and cost choices.

TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI
arXiv · 2026-07-24

TRACE-Router selects one model per agent task, keeps every call on that backend, and learns from the final outcome. Its benchmarks suggest task-level routing can improve accuracy and latency together.

Claude Opus 5 now available on AI Gateway
Vercel · 2026-07-24

AI Gateway now serves Claude Opus 5 with configurable reasoning, fast mode, fallbacks, and coding-agent setup; benign security tasks may still hit safeguards.

Notion's Token Town — Sarah Sachs, Notion
AI Engineer · 2026-07-23

Agent economics can regress even when token prices look stable. Route by task, preserve model optionality, and move deterministic work out of LLM calls before scaling usage.

AI Gateway now supports streaming transcription
Vercel · 2026-07-22

AI Gateway can now stream audio into transcription models and emit partial text, letting text-based agents accept lower-latency voice input without changing the agent itself.

Stable permalink · evidence auto-collected from source labels · synthesis maintained by feed7 editorial