# Inkling Small from Thinking Machines is now available on AI Gateway

Source: [Vercel](https://vercel.com/changelog/inkling-small-now-available-on-ai-gateway)  
Feed7 permalink: https://feed7.dev/p/inkling-small-now-available-on-ai-gateway-1a9781l  
Published: 2026-07-30T00:00:00.000Z  
Trust: Official Source (official_source)

## Why Included

Inkling Small is pitched as a lower-compute model for coding, tool use, and visual reasoning, with adjustable thinking effort and zero-data-retention routing through Vercel AI Gateway.

## Source Summary

Thinking Machines’ **Inkling Small** is available as **thinkingmachines/inkling-small**. Vercel says it is about **one quarter the size** of Inkling with comparable performance, supports audio and image reasoning, and offers adjustable thinking effort.

## Practical Implication

Test it as a cost-and-latency routing option for coding agents and tool-heavy tasks. Its programmatic crop, zoom, and inspection abilities may help when relevant details occupy a small part of a document or chart.

## Agent-Ready Context

Thinking Machines’ **Inkling Small** is available as **thinkingmachines/inkling-small**. Vercel says it is about **one quarter the size** of Inkling with comparable performance, supports audio and image reasoning, and offers adjustable thinking effort.

Test it as a cost-and-latency routing option for coding agents and tool-heavy tasks. Its programmatic crop, zoom, and inspection abilities may help when relevant details occupy a small part of a document or chart.

Comparable performance is not backed by benchmark figures in the supplied material. Zero Data Retention is supported, but must be enabled team-wide or per request and then depends on Gateway routing to eligible providers.

## Connected Context

Feed7 judgment across 297 accumulated Signals:

This introduces a smaller multimodal routing candidate for coding and tool-heavy work, with adjustable effort and document-detail inspection as differentiators. The quarter-size claim makes efficiency testing worthwhile, but absent comparative benchmarks it does not establish equivalent task quality, lower latency, or lower cost. Its ZDR value also depends on explicit configuration and eligible routing.

- [Why Large? Tiny LMs & Agents on Edge/Robotics — Cormac Brick, Google](https://feed7.dev/p/why-large-tiny-lms-agents-on-edge-robotics-cormac-brick-google-0fjif75) — The small-model guidance supports evaluating compact models where latency or resource constraints matter, while Inkling Small targets broader prompted coding and multimodal tool work rather than a narrow fine-tuned edge function.
- [GPT 5.6 Sol, Luna, and Terra now available on AI Gateway](https://feed7.dev/p/gpt-5-6-now-available-on-ai-gateway-106pgsr) — GPT-5.6 already offers tiered coding-agent routes; Inkling Small adds a separate compact multimodal candidate whose routing value still needs task-level measurement.
- [Claude Opus 5 now available on AI Gateway](https://feed7.dev/p/claude-opus-5-now-available-on-ai-gateway-16oaf27) — Both expose adjustable reasoning and visual capabilities, but Inkling Small is positioned as the compact option while Claude Opus 5 is the premium route, creating a concrete task-routing comparison.
- [Grok Voice Think Fast 2.0 now available on AI Gateway](https://feed7.dev/p/grok-voice-think-fast-2-0-now-available-on-ai-gateway-1dacr27) — Both combine multimodal reasoning with tool use, but Grok emphasizes realtime speech and early tool calls whereas Inkling Small emphasizes image inspection and compact deployment.

## Context Map

- Layer: model
- Domains: coding, image
- Topics: model-selection, reasoning, tool-use

## Uncertainty

- Comparable performance is not backed by benchmark figures in the supplied material. Zero Data Retention is supported, but must be enabled team-wide or per request and then depends on Gateway routing to eligible providers.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
