Sign InOpen Brain
VercelEngineering PostOfficial Source

DeepSeek overtakes Google on volume, cost per token falls 13.6%

Vercel’s July gateway data shows model routing, not list-price cuts, drove a 13.6% drop in average token cost as open-weight models gained production traffic.

Vercel · Aug 11, 2026
Open Source Open MarkdownOpen JSON
Source Summary

In Vercel AI Gateway’s July data, token volume rose **59%**, spend rose **37%**, and average price per token fell **13.6%**. DeepSeek reached **25% of token volume**, ahead of Google’s 10.7%.

Practical Implication

Builders should route agent tasks by capability tier and price, then keep measuring the mix. The report attributes the cost decline to teams shifting traffic toward cheaper models, not models becoming cheaper within a fixed mix.

Agent-Ready Context
In Vercel AI Gateway’s July data, token volume rose **59%**, spend rose **37%**, and average price per token fell **13.6%**. DeepSeek reached **25% of token volume**, ahead of Google’s 10.7%.

Builders should route agent tasks by capability tier and price, then keep measuring the mix. The report attributes the cost decline to teams shifting traffic toward cheaper models, not models becoming cheaper within a fixed mix.

This is traffic from one gateway, not a capability benchmark or a view of the whole market. Workload composition changed sharply, so volume share alone does not establish which model is best for coding agents.
Connected Context · Feed7 Judgment

This extends Vercel’s gateway evidence from a broad open-weight shift to a July mix dominated by DeepSeek, alongside lower blended token cost. It confirms that routing mix can materially change spend, while narrowing interpretation: the figures measure one gateway’s changing workloads, not model quality or a market-wide ranking, so task-level verification still governs selection.

Open-weight models surge to 29% of volume, price per token flattensThe June data established cheap open-weight volume alongside frontier agent workloads; the July Signal continues that routing pattern and identifies DeepSeek as the largest contributor to the newer mix.CFOs and the new economics of AICursor’s roughly ninefold cost variation across model families and widespread multi-model use reinforce the operational case for capability-tier routing behind the observed blended-cost decline.Open Source Is Dead. Long Live Open Source. — Saoud Rizwan, ClineCline’s verified-completion framing supplies the missing selection constraint: lower token prices are useful only when identical harness gates confirm that cheaper routed models complete the workload acceptably.Access and share AI Gateway leaderboard dataQueryable gateway data makes the Signal’s recommendation to keep measuring model mix implementable, while still providing usage evidence rather than a capability benchmark.
Context Map
industrycodingdata#model-selection#open-models#adoption
Uncertainty
This is traffic from one gateway, not a capability benchmark or a view of the whole market. Workload composition changed sharply, so volume share alone does not establish which model is best for coding agents.