Sign InOpen Brain
VercelEngineering PostOfficial Source

Ling 3.0 Tiny is now available on AI Gateway

Ling 3.0 Tiny gives coding agents a small MoE option with native function calling, prompt caching, a 256K context window, and gateway-based routing controls.

Vercel · Aug 6, 2026
Open Source Open MarkdownOpen JSON
Source Summary

**Ling 3.0 Tiny** is a mixture-of-experts model with **7.9B total parameters**, about **1.3B active per token**, a 256K-token context window, and up to 32K output tokens. It includes native function calling and prompt caching.

Practical Implication

Builders can select inclusionai/ling-3.0-tiny-free through the AI SDK or Vercel's coding-agent setup. Account for the scheduled switch to inclusionai/ling-3.0-tiny on August 14 when configuring persistent agent environments.

Agent-Ready Context
**Ling 3.0 Tiny** is a mixture-of-experts model with **7.9B total parameters**, about **1.3B active per token**, a 256K-token context window, and up to 32K output tokens. It includes native function calling and prompt caching.

Builders can select inclusionai/ling-3.0-tiny-free through the AI SDK or Vercel's coding-agent setup. Account for the scheduled switch to inclusionai/ling-3.0-tiny on August 14 when configuring persistent agent environments.

The free model slot ends at 8:00am PT on August 14 and replaces Ling 3.0 Flash. The material describes intended responsiveness but provides no coding evals or production reliability results.
Connected Context · Feed7 Judgment

This replaces Ling 3.0 Flash’s temporary evaluation slot with a smaller mixture-of-experts route that keeps 256K context while adding native function calling, prompt caching, and a scheduled model-ID transition. It narrows adoption to planned testing and configuration migration: parameter counts and intended responsiveness do not establish coding quality or production reliability, consistent with prior candidates’ case for workload-level evaluation.

Ling 3.0 Flash is now available on AI GatewayTiny directly replaces Flash in the free slot, preserving the 256K-context evaluation opportunity while requiring persistent environments to change model IDs on the stated date.Laguna S 2.1 is now available on AI GatewayLaguna provides a directly comparable 256K free coding route plus a paid 1M option, so Tiny’s context size alone does not distinguish it for selection.DeepSeek V4 Flash now runs updated weights on AI GatewayDeepSeek’s behavior changed behind an unchanged ID, while Ling requires an explicit scheduled ID switch; together they show that persistent routing needs both configuration tracking and repeated workload evaluation.AI Gateway adds unified fast mode supportGateway fast mode offers a separate latency control, so Ling’s intended responsiveness should not be treated as measured speed or as a substitute for verifying the serving tier and workload results.
Context Map
modelcoding#open-models#coding-agents#model-selection
Uncertainty
The free model slot ends at 8:00am PT on August 14 and replaces Ling 3.0 Flash. The material describes intended responsiveness but provides no coding evals or production reliability results.