# AI Gateway: GPT-5.6 pricing and speed updates

Source: [Vercel](https://vercel.com/changelog/ai-gateway-gpt-5-6-pricing-speed-updates)  
Feed7 permalink: https://feed7.dev/p/ai-gateway-gpt-5-6-pricing-speed-updates-06cxxee  
Published: 2026-07-30T00:00:00.000Z  
Trust: Official Source (official_source)

## Why Included

Vercel cut Luna and Terra token prices and raised Sol fast-mode speed without changing model IDs, so existing agent workloads inherit the changes without code edits.

## Source Summary

AI Gateway cut **GPT-5.6 Luna pricing by 80%** and **GPT-5.6 Terra pricing by 20%** across short and long contexts. Luna short-context rates are $0.20 input and $1.20 output per million tokens; Terra is $2 and $12.

## Practical Implication

Recheck model routing and budgets for agent workloads: Luna and Terra may now fit tasks previously assigned elsewhere. Existing requests receive the new rates because model IDs did not change.

## Agent-Ready Context

AI Gateway cut **GPT-5.6 Luna pricing by 80%** and **GPT-5.6 Terra pricing by 20%** across short and long contexts. Luna short-context rates are $0.20 input and $1.20 output per million tokens; Terra is $2 and $12.

Recheck model routing and budgets for agent workloads: Luna and Terra may now fit tasks previously assigned elsewhere. Existing requests receive the new rates because model IDs did not change.

**GPT-5.6 Sol fast mode is now 2.5x**, up from 1.5x, but its price is unchanged. The material provides only short-context dollar figures, so consult the full pricing detail before estimating long-context spend.

## Connected Context

Feed7 judgment across 297 accumulated Signals:

This materially changes the routing economics of existing GPT-5.6 workloads without requiring model-ID migrations: Luna becomes far cheaper, Terra moderately cheaper, and Sol fast mode faster at unchanged price. It justifies rerunning task-level cost and latency evaluations, but the missing long-context figures and outcome-quality evidence prevent price changes alone from determining the new routing policy.

- [AI Gateway adds unified fast mode support](https://feed7.dev/p/ai-gateway-adds-unified-fast-mode-support-144dq26) — Unified fast mode supplies the common request mechanism through which Sol’s increased fast-mode multiplier can affect existing coding-agent latency choices.
- [Notion's Token Town — Sarah Sachs, Notion](https://feed7.dev/p/notion-s-token-town-sarah-sachs-notion-1rb6gwh) — The price cuts strengthen the case for revisiting routes, while Token Town cautions that per-token savings alone do not establish lower end-to-end agent cost.
- [TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI](https://feed7.dev/p/2607-22465v1-1g7nw7j) — TRACE-Router provides an outcome-based method for testing whether the newly cheaper Luna or Terra should own an entire task instead of routing from price alone.
- [Routing rules now available on AI Gateway](https://feed7.dev/p/ai-gateway-routing-rules-024ifzc) — Routing rules make it possible to apply a reassessed Luna, Terra, or Sol policy centrally without changing application code, matching the operational consequence of the new economics.

## Context Map

- Layer: infra
- Domains: coding
- Topics: gateways, model-selection

## Uncertainty

- **GPT-5.6 Sol fast mode is now 2.5x**, up from 1.5x, but its price is unchanged. The material provides only short-context dollar figures, so consult the full pricing detail before estimating long-context spend.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
