Ling 3.0 Flash is now available on AI Gateway
Ling 3.0 Flash joins AI Gateway with a 256K context window, thinking and non-thinking modes, and free access through August 3 for agent workload testing.
AI Gateway now offers **Ling 3.0 Flash**, a Mixture-of-Experts model with **124B total and about 5.1B active parameters per token**. It has a **256K-token context window** and is free through **August 3**.
Builders can use the temporary free endpoint to evaluate high-frequency coding, document, and multi-step agent runs in both thinking and non-thinking modes, ideally against their own latency and token budgets.
AI Gateway now offers **Ling 3.0 Flash**, a Mixture-of-Experts model with **124B total and about 5.1B active parameters per token**. It has a **256K-token context window** and is free through **August 3**. Builders can use the temporary free endpoint to evaluate high-frequency coding, document, and multi-step agent runs in both thinking and non-thinking modes, ideally against their own latency and token budgets. The material provides architecture and positioning but no measured quality, speed, or reliability results. Free availability is time-limited, so avoid treating promotional pricing as the production cost baseline.
Ling 3.0 Flash adds another long-context coding-agent route to an increasingly crowded gateway model pool, but narrows immediate use to evaluation: unlike candidates with defined tiering or production-volume evidence, this signal supplies neither comparative results nor durable pricing. Its 256K context and thinking toggle most directly overlap Laguna S 2.1, making workload-specific measurement more useful than feature-list selection.