Ling 3.0 Tiny is now available on AI Gateway
Ling 3.0 Tiny gives coding agents a small MoE option with native function calling, prompt caching, a 256K context window, and gateway-based routing controls.
**Ling 3.0 Tiny** is a mixture-of-experts model with **7.9B total parameters**, about **1.3B active per token**, a 256K-token context window, and up to 32K output tokens. It includes native function calling and prompt caching.
Builders can select inclusionai/ling-3.0-tiny-free through the AI SDK or Vercel's coding-agent setup. Account for the scheduled switch to inclusionai/ling-3.0-tiny on August 14 when configuring persistent agent environments.
**Ling 3.0 Tiny** is a mixture-of-experts model with **7.9B total parameters**, about **1.3B active per token**, a 256K-token context window, and up to 32K output tokens. It includes native function calling and prompt caching. Builders can select inclusionai/ling-3.0-tiny-free through the AI SDK or Vercel's coding-agent setup. Account for the scheduled switch to inclusionai/ling-3.0-tiny on August 14 when configuring persistent agent environments. The free model slot ends at 8:00am PT on August 14 and replaces Ling 3.0 Flash. The material describes intended responsiveness but provides no coding evals or production reliability results.
This replaces Ling 3.0 Flash’s temporary evaluation slot with a smaller mixture-of-experts route that keeps 256K context while adding native function calling, prompt caching, and a scheduled model-ID transition. It narrows adoption to planned testing and configuration migration: parameter counts and intended responsiveness do not establish coding quality or production reliability, consistent with prior candidates’ case for workload-level evaluation.