{
  "schema_version": "1.1",
  "id": "s8:https://www.youtube.com/watch?v=_PdK6x7PQNM",
  "slug": "data-quality-is-the-compute-multiplier-ari-morcos-datologyai-0x7k2ve",
  "url": "https://feed7.dev/p/data-quality-is-the-compute-multiplier-ari-morcos-datologyai-0x7k2ve",
  "title": "Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI",
  "why_included": "Training-data curation can improve model quality and inference efficiency without simply adding compute. The practical work is decontamination, deduplication, balancing, task matching, and selective synthesis.",
  "summary": "Morcos presents data quality as a way to steepen model learning curves under constrained compute. Reported examples include multilingual performance above Qwen 3 with roughly **8× less compute**, similar performance to Qwen 3.5 with **35× fewer FLOPs per correct answer**, and about five points gained on LegalBench after **100B mid-training tokens**.",
  "practical_implication": "For model customization, optimize signal per token before buying a larger run: decontaminate benchmarks, remove semantic redundancy, balance topics, match the target task distribution, and synthesize variants from selected high-quality documents. Sequence mixtures deliberately across training phases.",
  "agent_context": "Morcos presents data quality as a way to steepen model learning curves under constrained compute. Reported examples include multilingual performance above Qwen 3 with roughly **8× less compute**, similar performance to Qwen 3.5 with **35× fewer FLOPs per correct answer**, and about five points gained on LegalBench after **100B mid-training tokens**.\n\nFor model customization, optimize signal per token before buying a larger run: decontaminate benchmarks, remove semantic redundancy, balance topics, match the target task distribution, and synthesize variants from selected high-quality documents. Sequence mixtures deliberately across training phases.\n\nThese are selected results from a data-curation vendor's talk, with limited experimental detail in the transcript. The gains span different models, tasks, and measurement families, so they should not be read as one universal multiplier or expected outcome.",
  "source": {
    "name": "AI Engineer",
    "url": "https://www.youtube.com/watch?v=_PdK6x7PQNM",
    "published_at": "2026-07-31T23:00:06.000Z"
  },
  "source_class": "video",
  "content_type": "Video",
  "layer": "model",
  "domains": [
    "data"
  ],
  "topics": [
    "model-selection",
    "open-models",
    "reasoning"
  ],
  "verification": {
    "status": "source_linked",
    "label": "Source Linked",
    "method": "source_feed",
    "verified_at": null
  },
  "uncertainty": [
    "These are selected results from a data-curation vendor's talk, with limited experimental detail in the transcript. The gains span different models, tasks, and measurement families, so they should not be read as one universal multiplier or expected outcome."
  ],
  "connected_context": {
    "meaning": "This shifts efficiency analysis upstream: model capability per unit of compute may depend as much on selecting, balancing, sequencing, and decontaminating data as on choosing a larger model or training objective. It supports targeted data mixtures and synthetic variants, but the vendor-selected results do not justify treating the reported gains as a portable multiplier across models, tasks, or deployments.",
    "corpus_size": 318,
    "generated_at": "2026-08-01T10:08:32.212Z",
    "connections": [
      {
        "title": "The Base Model Is Dead — Varun Singh, Arcee AI",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=xbPriQWXtWM",
        "feed7_url": "https://feed7.dev/p/the-base-model-is-dead-varun-singh-arcee-ai-02hts76",
        "reason": "The base-model data shift toward code, reasoning, and agent priors supports task-matched mixtures; Morcos adds concrete curation and phase-sequencing practices to that unresolved data-design choice."
      },
      {
        "title": "Introducing Grok 4.5",
        "source_name": "Cursor",
        "source_url": "https://cursor.com/blog/grok-4-5",
        "feed7_url": "https://feed7.dev/p/grok-4-5-1n0zgxx",
        "reason": "The excluded contaminated CursorBench result illustrates why benchmark decontamination is a prerequisite for interpreting apparent gains from training data."
      },
      {
        "title": "Program-as-Weights: A Programming Paradigm for Fuzzy Functions",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2607.02512v1",
        "feed7_url": "https://feed7.dev/p/2607-02512v1-1dr5458",
        "reason": "Program-as-Weights offers a complementary efficiency route—small task adapters rather than curated large-scale training—so both should be evaluated as alternatives to simply buying a larger run."
      },
      {
        "title": "$β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2607.28582v1",
        "feed7_url": "https://feed7.dev/p/2607-28582v1-0egi1xh",
        "reason": "β-OPSD changes the optimization objective and teacher-reference balance, contrasting with this talk’s claim that better signal selection and sequencing can improve learning before changing the training algorithm."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-07-31T23:00:06.000Z",
  "modified_at": "2026-07-31T23:00:06.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/data-quality-is-the-compute-multiplier-ari-morcos-datologyai-0x7k2ve",
    "json": "https://feed7.dev/p/data-quality-is-the-compute-multiplier-ari-morcos-datologyai-0x7k2ve.json",
    "markdown": "https://feed7.dev/p/data-quality-is-the-compute-multiplier-ari-morcos-datologyai-0x7k2ve.md"
  }
}