Kimi K3 and Kimi K3 Fast with ZDR and US-based providers now on AI Gateway
Kimi K3 now has US-hosted, ZDR-capable gateway routes plus a faster tier, giving coding-agent users explicit latency, residency, retention, and cost choices.
AI Gateway now serves **Kimi K3 and Kimi K3 Fast** through US-based providers including Baseten and Fireworks. Both support **Zero Data Retention**, while the gateway can route across providers for fallback and capacity.
For coding agents, use moonshotai/kimi-k3 and select the speed option when latency matters; it falls back to standard speed if the fast tier is unavailable. Pin inferenceRegion for US-only processing, and inspect the endpoints API before choosing a provider or variant.
AI Gateway now serves **Kimi K3 and Kimi K3 Fast** through US-based providers including Baseten and Fireworks. Both support **Zero Data Retention**, while the gateway can route across providers for fallback and capacity. For coding agents, use moonshotai/kimi-k3 and select the speed option when latency matters; it falls back to standard speed if the fast tier is unavailable. Pin inferenceRegion for US-only processing, and inspect the endpoints API before choosing a provider or variant. Kimi K3 Fast costs **about 50% more** than the base model, while US regional inference is **about 10% more** than the regular variant. Exact prices and capabilities vary by provider.
This expands coding-agent routing with operationally differentiated Kimi endpoints: standard versus faster service, US-only processing, ZDR, and provider fallback. It clarifies concrete latency, residency, retention, and price knobs, but supplies no quality comparison, so the new routes belong in workload testing rather than replacing existing models on specifications alone.