Sign InOpen Brain
VercelEngineering PostOfficial Source

GLM 5.3 FlashX now available on AI Gateway

GLM 5.3 FlashX brings roughly 200 tokens-per-second serving to Vercel AI Gateway, offering a lower-wait option for coding-agent output and repeated tool loops.

Vercel · Sep 18, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Vercel AI Gateway now serves **GLM 5.3 FlashX** under `zai/glm-5.3-flashx`, with reported throughput of **about 200 tokens per second** for the multimodal coding model.

Practical Implication

For latency-sensitive agents, test it on interactive output and tool-heavy loops. Gateway setup also provides usage and cost tracking, retries, failover, budgets, and routing through one API.

Agent-Ready Context
Vercel AI Gateway now serves **GLM 5.3 FlashX** under `zai/glm-5.3-flashx`, with reported throughput of **about 200 tokens per second** for the multimodal coding model.

For latency-sensitive agents, test it on interactive output and tool-heavy loops. Gateway setup also provides usage and cost tracking, retries, failover, budgets, and routing through one API.

The material reports serving speed, not task quality, time to first token, or end-to-end tool-loop latency. Provider pricing has no gateway markup, but actual workload cost is not given.
Connected Context · Feed7 Judgment

This adds a high-throughput multimodal coding candidate to an increasingly crowded gateway catalog, making interactive and tool-heavy trials more plausible. The reported token rate narrows only serving-speed expectations; selection still requires workload tests of first-token delay, task quality, full loop time, reliability, and total cost.

DeepSeek V4.1 Flash now available on AI GatewayDeepSeek V4.1 Flash is a directly comparable multimodal coding route, reinforcing that feature coverage and capacity claims require workload-specific comparison.Qwen 3.8 Max 0902 now available on AI GatewayThe pinned Qwen snapshot highlights a reproducibility property absent from the GLM announcement and useful when evaluating model behavior over time.Kimi K3 and Kimi K3 Fast with ZDR and US-based providers now on AI GatewayKimi’s differentiated hosting, retention, speed, and price options show operational dimensions that the GLM throughput figure alone does not resolve.Claude Opus 5 now available on AI GatewayClaude Opus 5 provides another speed-configurable coding route, making quality, safeguards, and end-to-end behavior—not raw throughput alone—the relevant routing comparison.
Context Map
toolscoding#coding-agents#gateways#model-selection
Uncertainty
The material reports serving speed, not task quality, time to first token, or end-to-end tool-loop latency. Provider pricing has no gateway markup, but actual workload cost is not given.