# The Base Model Is Dead — Varun Singh, Arcee AI

Source: [AI Engineer](https://www.youtube.com/watch?v=xbPriQWXtWM)  
Feed7 permalink: https://feed7.dev/p/the-base-model-is-dead-varun-singh-arcee-ai-02hts76  
Published: 2026-07-31T20:30:21.000Z  
Trust: Source Linked (source_linked)

## Why Included

Base-model data is shifting from broad web imitation toward code, reasoning, and agent-task priors. The unresolved choice is how early to introduce synthetic and instruction-shaped data.

## Source Summary

The talk contrasts GPT-3’s roughly **85% web-derived mix** with MAI Thinking 1 at **15% web text**. Newer recipes emphasize code, STEM, reasoning traces, and task-shaped data that better prepare models for downstream RL.

## Practical Implication

When selecting or training a model for agents, evaluate its pre-RL skill coverage, not just general knowledge. **NeMoTron 3 Ultra** pulls SFT-style data into pre-training, while synthetic rephrasing can expose the same information in several forms.

## Agent-Ready Context

The talk contrasts GPT-3’s roughly **85% web-derived mix** with MAI Thinking 1 at **15% web text**. Newer recipes emphasize code, STEM, reasoning traces, and task-shaped data that better prepare models for downstream RL.

When selecting or training a model for agents, evaluate its pre-RL skill coverage, not just general knowledge. **NeMoTron 3 Ultra** pulls SFT-style data into pre-training, while synthetic rephrasing can expose the same information in several forms.

There is no settled recipe: MAI Thinking 1 deliberately avoids synthetic model-generated data, while NeMoTron leans into it. Synthetic data can degrade a model when used indiscriminately, and it remains unclear how far RL can displace supervised learning for language.

## Connected Context

Feed7 judgment across 318 accumulated Signals:

This shifts model selection beneath post-training results: agent readiness may depend on whether pre-training already covers code, STEM, reasoning, and task-shaped behavior. It makes gateway tiers, context sizes, and reasoning toggles insufficient selection criteria on their own. The conflicting MAI and NeMoTron recipes also prevent a general rule about synthetic data, so provenance and workload evaluation matter more than adopting either synthetic-heavy or synthetic-free training as doctrine.

- [Introducing Grok 4.5](https://feed7.dev/p/grok-4-5-1n0zgxx) — Grok 4.5’s benchmark exclusion because training included an earlier code snapshot reinforces this signal’s demand to inspect pre-training composition and provenance, not interpret downstream scores without qualification.
- [GPT 5.6 Sol, Luna, and Terra now available on AI Gateway](https://feed7.dev/p/gpt-5-6-now-available-on-ai-gateway-106pgsr) — Gateway tiering offers convenient flagship, balanced, and lower-cost routes, but this signal says those labels must be supplemented by testing whether each model’s pre-RL skill coverage matches the agent workload.
- [DeepSeek V4 Flash now runs updated weights on AI Gateway](https://feed7.dev/p/deepseek-v4-flash-now-runs-updated-weights-on-ai-gateway-1qqe8mw) — The reported score jump after a weight replacement shows that underlying training changes can materially alter a fixed model route, while this signal cautions that one downstream benchmark does not reveal the broader data recipe or skill coverage.
- [What's Next After RLHF? — Diogo Almeida, TypeSafe AI](https://feed7.dev/p/what-s-next-after-rlhf-diogo-almeida-typesafe-ai-1scytnx) — The RLHF critique supports this signal’s unresolved boundary between supervised learning and RL: post-training for persuasive interaction does not establish the dependable decision-making skills needed for autonomous agents.

## Context Map

- Layer: model
- Domains: coding
- Topics: reasoning, coding-agents, model-selection

## Uncertainty

- There is no settled recipe: MAI Thinking 1 deliberately avoids synthetic model-generated data, while NeMoTron leans into it. Synthetic data can degrade a model when used indiscriminately, and it remains unclear how far RL can displace supervised learning for language.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
