GLM 5.3 FlashX now available on AI Gateway
GLM 5.3 FlashX brings roughly 200 tokens-per-second serving to Vercel AI Gateway, offering a lower-wait option for coding-agent output and repeated tool loops.
Vercel AI Gateway now serves **GLM 5.3 FlashX** under `zai/glm-5.3-flashx`, with reported throughput of **about 200 tokens per second** for the multimodal coding model.
For latency-sensitive agents, test it on interactive output and tool-heavy loops. Gateway setup also provides usage and cost tracking, retries, failover, budgets, and routing through one API.
Vercel AI Gateway now serves **GLM 5.3 FlashX** under `zai/glm-5.3-flashx`, with reported throughput of **about 200 tokens per second** for the multimodal coding model. For latency-sensitive agents, test it on interactive output and tool-heavy loops. Gateway setup also provides usage and cost tracking, retries, failover, budgets, and routing through one API. The material reports serving speed, not task quality, time to first token, or end-to-end tool-loop latency. Provider pricing has no gateway markup, but actual workload cost is not given.
This adds a high-throughput multimodal coding candidate to an increasingly crowded gateway catalog, making interactive and tool-heavy trials more plausible. The reported token rate narrows only serving-speed expectations; selection still requires workload tests of first-token delay, task quality, full loop time, reliability, and total cost.