# GLM 5.3 FlashX now available on AI Gateway

Source: [Vercel](https://vercel.com/changelog/glm-5-3-flashx-now-available-on-ai-gateway)  
Feed7 permalink: https://feed7.dev/p/glm-5-3-flashx-now-available-on-ai-gateway-1yir9b2  
Published: 2026-09-18T00:00:00.000Z  
Trust: Official Source (official_source)

## Why Included

GLM 5.3 FlashX brings roughly 200 tokens-per-second serving to Vercel AI Gateway, offering a lower-wait option for coding-agent output and repeated tool loops.

## Source Summary

Vercel AI Gateway now serves **GLM 5.3 FlashX** under `zai/glm-5.3-flashx`, with reported throughput of **about 200 tokens per second** for the multimodal coding model.

## Practical Implication

For latency-sensitive agents, test it on interactive output and tool-heavy loops. Gateway setup also provides usage and cost tracking, retries, failover, budgets, and routing through one API.

## Agent-Ready Context

Vercel AI Gateway now serves **GLM 5.3 FlashX** under `zai/glm-5.3-flashx`, with reported throughput of **about 200 tokens per second** for the multimodal coding model.

For latency-sensitive agents, test it on interactive output and tool-heavy loops. Gateway setup also provides usage and cost tracking, retries, failover, budgets, and routing through one API.

The material reports serving speed, not task quality, time to first token, or end-to-end tool-loop latency. Provider pricing has no gateway markup, but actual workload cost is not given.

## Connected Context

Feed7 judgment across 807 accumulated Signals:

This adds a high-throughput multimodal coding candidate to an increasingly crowded gateway catalog, making interactive and tool-heavy trials more plausible. The reported token rate narrows only serving-speed expectations; selection still requires workload tests of first-token delay, task quality, full loop time, reliability, and total cost.

- [DeepSeek V4.1 Flash now available on AI Gateway](https://feed7.dev/p/deepseek-v4-1-flash-now-available-on-ai-gateway-11mmjw1) — DeepSeek V4.1 Flash is a directly comparable multimodal coding route, reinforcing that feature coverage and capacity claims require workload-specific comparison.
- [Qwen 3.8 Max 0902 now available on AI Gateway](https://feed7.dev/p/qwen-3-8-max-0902-now-available-on-ai-gateway-0zxju42) — The pinned Qwen snapshot highlights a reproducibility property absent from the GLM announcement and useful when evaluating model behavior over time.
- [Kimi K3 and Kimi K3 Fast with ZDR and US-based providers now on AI Gateway](https://feed7.dev/p/kimi-k3-and-kimi-k3-fast-on-ai-gateway-0jpdaz6) — Kimi’s differentiated hosting, retention, speed, and price options show operational dimensions that the GLM throughput figure alone does not resolve.
- [Claude Opus 5 now available on AI Gateway](https://feed7.dev/p/claude-opus-5-now-available-on-ai-gateway-16oaf27) — Claude Opus 5 provides another speed-configurable coding route, making quality, safeguards, and end-to-end behavior—not raw throughput alone—the relevant routing comparison.

## Context Map

- Layer: tools
- Domains: coding
- Topics: coding-agents, gateways, model-selection

## Uncertainty

- The material reports serving speed, not task quality, time to first token, or end-to-end tool-loop latency. Provider pricing has no gateway markup, but actual workload cost is not given.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
