Inkling Small from Thinking Machines is now available on AI Gateway
Inkling Small is pitched as a lower-compute model for coding, tool use, and visual reasoning, with adjustable thinking effort and zero-data-retention routing through Vercel AI Gateway.
Thinking Machines’ **Inkling Small** is available as **thinkingmachines/inkling-small**. Vercel says it is about **one quarter the size** of Inkling with comparable performance, supports audio and image reasoning, and offers adjustable thinking effort.
Test it as a cost-and-latency routing option for coding agents and tool-heavy tasks. Its programmatic crop, zoom, and inspection abilities may help when relevant details occupy a small part of a document or chart.
Thinking Machines’ **Inkling Small** is available as **thinkingmachines/inkling-small**. Vercel says it is about **one quarter the size** of Inkling with comparable performance, supports audio and image reasoning, and offers adjustable thinking effort. Test it as a cost-and-latency routing option for coding agents and tool-heavy tasks. Its programmatic crop, zoom, and inspection abilities may help when relevant details occupy a small part of a document or chart. Comparable performance is not backed by benchmark figures in the supplied material. Zero Data Retention is supported, but must be enabled team-wide or per request and then depends on Gateway routing to eligible providers.
This introduces a smaller multimodal routing candidate for coding and tool-heavy work, with adjustable effort and document-detail inspection as differentiators. The quarter-size claim makes efficiency testing worthwhile, but absent comparative benchmarks it does not establish equivalent task quality, lower latency, or lower cost. Its ZDR value also depends on explicit configuration and eligible routing.