{
  "schema_version": "1.1",
  "id": "archive:https://arxiv.org/abs/2608.20290v1",
  "slug": "2608-20290v1-13o5gau",
  "url": "https://feed7.dev/p/2608-20290v1-13o5gau",
  "title": "Phantom Gains: Auditing Self-Improvement Against a Measured Null",
  "why_included": "Per-problem self-improvement claims can arise from inference and evaluation noise. This audit argues every transition statistic needs a measured null from frozen baseline replicates.",
  "summary": "The study ran three rounds of rank-32 LoRA self-training on Qwen3-8B alongside a frozen control using the same pipeline. It found **seven measurement failures** capable of reversing findings when the control was omitted.",
  "practical_implication": "For agent and model evaluations, do not treat problem-level gains and losses as ground truth from one decode. Measure each statistic’s null with baseline replicates, then use per-problem tests and false-discovery-rate control.",
  "agent_context": "The study ran three rounds of rank-32 LoRA self-training on Qwen3-8B alongside a frozen control using the same pipeline. It found **seven measurement failures** capable of reversing findings when the control was omitted.\n\nFor agent and model evaluations, do not treat problem-level gains and losses as ground truth from one decode. Measure each statistic’s null with baseline replicates, then use per-problem tests and false-discovery-rate control.\n\nA single greedy decode gave an untrained model an apparent expansion rate of **0.280**, while the replacement exact test detected **nothing on held-out replicates**. Results for problems the base model never solved remained inconclusive, and the method needs more baseline replicates than many studies possess.",
  "source": {
    "name": "arXiv",
    "url": "https://arxiv.org/abs/2608.20290v1",
    "published_at": "2026-08-20T17:30:14.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Paper",
  "layer": "benchmark",
  "domains": [
    "research"
  ],
  "topics": [
    "benchmark-integrity",
    "agent-evals",
    "agent-reliability"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "A single greedy decode gave an untrained model an apparent expansion rate of **0.280**, while the replacement exact test detected **nothing on held-out replicates**. Results for problems the base model never solved remained inconclusive, and the method needs more baseline replicates than many studies possess."
  ],
  "connected_context": {
    "meaning": "Phantom Gains converts general benchmark skepticism into a concrete prerequisite for self-improvement claims: estimate a statistic’s null from frozen-control replicates before interpreting per-problem transitions. It shows that one-decode expansion can vanish under a replacement exact test and FDR control. Repeated sampling alone is therefore insufficient unless baseline variability is measured; unsolved-base cases remain unresolved.",
    "corpus_size": 525,
    "generated_at": "2026-08-22T21:15:25.320Z",
    "connections": [
      {
        "title": "Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2607.28576v1",
        "feed7_url": "https://feed7.dev/p/2607-28576v1-09h2m1u",
        "reason": "The repeated-sampling study establishes a token-matched baseline for reflection claims; Phantom Gains adds that repeated decodes must also estimate frozen-baseline variability and control multiplicity before apparent per-problem changes count as improvement."
      },
      {
        "title": "When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=-npY6XjM8CQ",
        "feed7_url": "https://feed7.dev/p/when-will-the-benchmaxxing-plague-end-nick-heiner-surge-ai-178gqcg",
        "reason": "The benchmaxxing critique broadly warns that leaderboard gains can be artifacts; Phantom Gains supplies a specific statistical mechanism and control protocol for detecting such artifacts in self-training evaluations."
      },
      {
        "title": "SocietyBench: Forecasting Counterfactual Social-World Evolution",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2608.04009v1",
        "feed7_url": "https://feed7.dev/p/2608-04009v1-05m5u8w",
        "reason": "SocietyBench warns against generalizing variable aggregate results and recommends event-level reporting; Phantom Gains shows that problem-level reporting itself requires a measured null and false-discovery control to avoid mistaking decode noise for transitions."
      },
      {
        "title": "Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2608.07437v1",
        "feed7_url": "https://feed7.dev/p/2608-07437v1-0t4v96m",
        "reason": "Fisher-R1’s reported single-trial improvement is a model-comparison result, while Phantom Gains shows what additional frozen-control replication and per-problem testing would be required before interpreting transition-level changes as reliable self-improvement."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-08-20T17:30:14.000Z",
  "modified_at": "2026-08-20T17:30:14.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/2608-20290v1-13o5gau",
    "json": "https://feed7.dev/p/2608-20290v1-13o5gau.json",
    "markdown": "https://feed7.dev/p/2608-20290v1-13o5gau.md"
  }
}