# DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

Source: [arXiv](https://arxiv.org/abs/2607.24717v1)  
Feed7 permalink: https://feed7.dev/p/2607-24717v1-0468y8t  
Published: 2026-07-27T17:54:12.000Z  
Trust: Needs Review (needs_review)

## Why Included

DataOrchestra chooses a processing pipeline per pre-training example, improving average benchmark results while avoiding compute on chunks that need no transformation.

## Source Summary

DataOrchestra routes each data chunk to drop, preserve, or clean, then chooses programmatic edits or LLM rewrites and generates rewrite instructions. Models from **0.5B to 7B** trained on its output showed stable average gains across **11 benchmarks**.

## Practical Implication

Builders preparing training corpora should reconsider corpus-wide cleaning rules. Per-example routing can reserve expensive rewriting for material that needs it, and the reported math continued-pretraining results also exceeded stronger processing baselines.

## Agent-Ready Context

DataOrchestra routes each data chunk to drop, preserve, or clean, then chooses programmatic edits or LLM rewrites and generates rewrite instructions. Models from **0.5B to 7B** trained on its output showed stable average gains across **11 benchmarks**.

Builders preparing training corpora should reconsider corpus-wide cleaning rules. Per-example routing can reserve expensive rewriting for material that needs it, and the reported math continued-pretraining results also exceeded stronger processing baselines.

The abstract gives no gain sizes, processing-cost figures, or operational complexity. It therefore supports the routing principle more clearly than any estimate of whether implementing the orchestrator pays off for a particular dataset.

## Context Map

- Layer: model
- Domains: data
- Topics: open-models

## Uncertainty

- The abstract gives no gain sizes, processing-cost figures, or operational complexity. It therefore supports the routing principle more clearly than any estimate of whether implementing the orchestrator pays off for a particular dataset.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
