# DeepSeek V4 Flash now runs updated weights on AI Gateway

Source: [Vercel](https://vercel.com/changelog/deepseek-v4-flash-now-runs-updated-weights-on-ai-gateway)  
Feed7 permalink: https://feed7.dev/p/deepseek-v4-flash-now-runs-updated-weights-on-ai-gateway-1qqe8mw  
Published: 2026-07-31T07:00:00.000Z  
Trust: Official Source (official_source)

## Why Included

DeepSeek V4 Flash’s updated weights replace the preview behind the existing model ID, raising its reported Terminal-Bench score from 56.9 to 82.7 without code changes.

## Source Summary

Updated **DeepSeek V4 Flash** weights now load automatically for the existing AI Gateway model ID. Vercel reports a **Terminal-Bench score of 82.7**, versus **56.9** for the April preview.

## Practical Implication

Re-evaluate the model on your own coding-agent tasks before changing routing. Existing integrations need no code or model-ID change, so behavior may shift even if your configuration stays fixed.

## Agent-Ready Context

Updated **DeepSeek V4 Flash** weights now load automatically for the existing AI Gateway model ID. Vercel reports a **Terminal-Bench score of 82.7**, versus **56.9** for the April preview.

Re-evaluate the model on your own coding-agent tasks before changing routing. Existing integrations need no code or model-ID change, so behavior may shift even if your configuration stays fixed.

Only DeepSeek serves the updated weights for now. Other providers, including Zero Data Retention options, were announced for the following week but are not yet part of this release.

## Connected Context

Feed7 judgment across 318 accumulated Signals:

This turns model selection into a versioning concern: an existing route can change materially without a model-ID or code change. The higher reported benchmark supports renewed evaluation but does not establish superiority on a team’s own tasks. Compared with newly added gateway models, this is inherited behavioral drift, and the single-provider rollout temporarily limits routing and Zero Data Retention choices.

- [ReviewDebt: a practical framework for scoring every pull request — Sachin Gupta, Ebay](https://feed7.dev/p/reviewdebt-a-practical-framework-for-scoring-every-pull-request-sachin-g-0iyjtyk) — The model benchmark measures task performance, while ReviewDebt measures downstream verification burden; together they imply that a stronger coding score alone does not determine whether the route improves an engineering workflow.

## Context Map

- Layer: model
- Domains: coding
- Topics: model-selection, coding-agents, agent-evals

## Uncertainty

- Only DeepSeek serves the updated weights for now. Other providers, including Zero Data Retention options, were announced for the following week but are not yet part of this release.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
