{
  "schema_version": "1.1",
  "id": "archive:https://openai.com/index/previewing-ultrafast",
  "slug": "previewing-ultrafast-07sphgw",
  "url": "https://feed7.dev/p/previewing-ultrafast-07sphgw",
  "title": "Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed",
  "why_included": "OpenAI’s preview tier runs GPT-5.6 Sol at up to 750 output tokens per second, targeting agent workloads where generation latency is the bottleneck.",
  "summary": "OpenAI is previewing **Ultrafast**, an API service tier for **GPT-5.6 Sol** powered by Cerebras. It claims up to **14× faster** operation and up to **750 output tokens per second**.",
  "practical_implication": "Builders with latency-bound coding agents should test whether faster generation materially shortens end-to-end runs, especially when tool execution is already quick.",
  "agent_context": "OpenAI is previewing **Ultrafast**, an API service tier for **GPT-5.6 Sol** powered by Cerebras. It claims up to **14× faster** operation and up to **750 output tokens per second**.\n\nBuilders with latency-bound coding agents should test whether faster generation materially shortens end-to-end runs, especially when tool execution is already quick.\n\nBoth figures are qualified maxima. The material gives no pricing, availability, workload methodology, or comparison baseline, and tool latency may still dominate many runs.",
  "source": {
    "name": "OpenAI",
    "url": "https://openai.com/index/previewing-ultrafast",
    "published_at": "2026-08-13T10:00:00.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Official Release",
  "layer": "infra",
  "domains": [
    "coding"
  ],
  "topics": [
    "coding-agents"
  ],
  "verification": {
    "status": "official_source",
    "label": "Official Source",
    "method": "source_feed",
    "verified_at": null
  },
  "uncertainty": [
    "Both figures are qualified maxima. The material gives no pricing, availability, workload methodology, or comparison baseline, and tool latency may still dominate many runs."
  ],
  "connected_context": {
    "meaning": "Ultrafast adds a model-specific serving option whose claimed ceiling is substantially more concrete than the gateway’s generic fast-mode control, but it does not establish faster completed agent work. It narrows evaluation to latency-bound runs where generation is material, with tool time, price, access, and workload-level outcomes still unresolved.",
    "corpus_size": 479,
    "generated_at": "2026-08-18T10:04:07.972Z",
    "connections": [
      {
        "title": "AI Gateway adds unified fast mode support",
        "source_name": "Vercel",
        "source_url": "https://vercel.com/changelog/ai-gateway-adds-unified-fast-mode-support",
        "feed7_url": "https://feed7.dev/p/ai-gateway-adds-unified-fast-mode-support-144dq26",
        "reason": "The gateway candidate provides a common client-side fast-mode control, while Ultrafast describes a specific GPT-5.6 Sol service tier; together they separate routing convenience from evidence of actual speed on a given route."
      },
      {
        "title": "GPT-5.6 Sol is 50% off on AI Gateway for the next month",
        "source_name": "Vercel",
        "source_url": "https://vercel.com/changelog/gpt-5-6-sol-is-50-off-on-ai-gateway-for-the-next-month",
        "feed7_url": "https://feed7.dev/p/gpt-5-6-sol-is-50-off-on-ai-gateway-for-the-next-month-1i67zht",
        "reason": "The discount supplies a temporary cost-testing window for the same model, but the texts do not establish that discounted gateway requests use Ultrafast, so price and speed must be evaluated as separate variables."
      },
      {
        "title": "How Forward Deployed Engineering is done at Cognition — Jia Wu",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=RVxym6mmIns",
        "feed7_url": "https://feed7.dev/p/how-forward-deployed-engineering-is-done-at-cognition-jia-wu-06h8ybj",
        "reason": "Cognition’s outcome-based measurement standard supplies the missing evaluation frame: higher token throughput matters only if it shortens delivery or increases accepted work."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-08-13T10:00:00.000Z",
  "modified_at": "2026-08-13T10:00:00.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/previewing-ultrafast-07sphgw",
    "json": "https://feed7.dev/p/previewing-ultrafast-07sphgw.json",
    "markdown": "https://feed7.dev/p/previewing-ultrafast-07sphgw.md"
  }
}