{
  "schema_version": "1.1",
  "id": "archive:https://arxiv.org/abs/2608.13515v1",
  "slug": "2608-13515v1-1cot0y6",
  "url": "https://feed7.dev/p/2608-13515v1-1cot0y6",
  "title": "Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining",
  "why_included": "A task-agnostic influence measure tracks which examples steer pretraining toward final parameters without choosing a downstream eval. It reveals a literature-to-STEM shift across training.",
  "summary": "The method defines an example’s influence by how much its gradient update reduces squared distance to the run’s final parameters. It estimates that value from intermediate checkpoints **without retraining** or selecting a downstream task.",
  "practical_implication": "Across **18 Pythia and PolyPythia configurations**, influential data changed over time: literature-related examples aligned more strongly early, while STEM data aligned more strongly later. Training-pipeline builders can use this trajectory view alongside task-specific attribution when inspecting data mixtures.",
  "agent_context": "The method defines an example’s influence by how much its gradient update reduces squared distance to the run’s final parameters. It estimates that value from intermediate checkpoints **without retraining** or selecting a downstream task.\n\nAcross **18 Pythia and PolyPythia configurations**, influential data changed over time: literature-related examples aligned more strongly early, while STEM data aligned more strongly later. Training-pipeline builders can use this trajectory view alongside task-specific attribution when inspecting data mixtures.\n\nThe crossover was qualitative and broadly consistent, but the measure is tied to the final parameters of a particular run. It does not show which examples improve a coding agent, nor replace downstream capability and reliability evaluations.",
  "source": {
    "name": "arXiv",
    "url": "https://arxiv.org/abs/2608.13515v1",
    "published_at": "2026-08-13T17:36:49.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Paper",
  "layer": "benchmark",
  "domains": [
    "research",
    "data"
  ],
  "topics": [
    "benchmark-integrity"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "The crossover was qualitative and broadly consistent, but the measure is tied to the final parameters of a particular run. It does not show which examples improve a coding agent, nor replace downstream capability and reliability evaluations."
  ],
  "connected_context": {
    "meaning": "This adds a task-agnostic, checkpoint-based view of which pretraining examples most shaped a particular run, complementing candidates focused on downstream behavior and benchmark design. The changing literature-to-STEM influence pattern argues against treating data contribution as static, but the run-relative measure does not establish agent capability or reliability and therefore remains a diagnostic alongside, not a substitute for, task and workflow evaluations.",
    "corpus_size": 462,
    "generated_at": "2026-08-16T10:04:38.518Z",
    "connections": [
      {
        "title": "LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2608.13545v1",
        "feed7_url": "https://feed7.dev/p/2608-13545v1-1ray1wb",
        "reason": "LittleLearner controls prior exposure to study knowledge acquisition, while this method retrospectively estimates how individual examples influenced final parameters; together they offer complementary experimental and diagnostic views of learning from data."
      },
      {
        "title": "State of Data — Sean Cai, Independent / State of Data",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=ZyIoTOAbRfs",
        "feed7_url": "https://feed7.dev/p/state-of-data-sean-cai-independent-state-of-data-0v9fy69",
        "reason": "The workflow-trace candidate emphasizes scaffold-dependent downstream behavior, whereas this Signal measures pretraining influence without choosing a downstream task; the two operate at different layers and neither substitutes for the other."
      },
      {
        "title": "LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2607.02513v1",
        "feed7_url": "https://feed7.dev/p/2607-02513v1-0lwaytn",
        "reason": "LACUNA supplies ground-truth weight locations for injected information to test unlearning, while this Signal estimates example influence relative to final parameters without such ground truth; LACUNA therefore highlights a validation advantage absent from this broader attribution method."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-08-13T17:36:49.000Z",
  "modified_at": "2026-08-13T17:36:49.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/2608-13515v1-1cot0y6",
    "json": "https://feed7.dev/p/2608-13515v1-1cot0y6.json",
    "markdown": "https://feed7.dev/p/2608-13515v1-1cot0y6.md"
  }
}