{
  "schema_version": "1.1",
  "id": "s13:https://arxiv.org/abs/2607.28576v1",
  "slug": "2607-28576v1-09h2m1u",
  "url": "https://feed7.dev/p/2607-28576v1-09h2m1u",
  "title": "Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B",
  "why_included": "At equal generated-token cost, repeated sampling matched or beat seven reflection, critique, selection, and debate methods. Agent evals should budget every generated token, not compare against one-shot baselines.",
  "summary": "Researchers tested **seven methods** on 1.5B, 3B, and 7B open models across two math benchmarks, counting critique, reflection, debate, and checking tokens. Across **36 paired comparisons**, none reliably beat repeated sampling at equal cost.",
  "practical_implication": "For agent workflows, benchmark reflection loops against repeated independent attempts using the same measured token budget. Self-inspection deserves particular scrutiny: **all 18 comparisons were negative**, and ten results were reliably worse.",
  "agent_context": "Researchers tested **seven methods** on 1.5B, 3B, and 7B open models across two math benchmarks, counting critique, reflection, debate, and checking tokens. Across **36 paired comparisons**, none reliably beat repeated sampling at equal cost.\n\nFor agent workflows, benchmark reflection loops against repeated independent attempts using the same measured token budget. Self-inspection deserves particular scrutiny: **all 18 comparisons were negative**, and ten results were reliably worse.\n\nThe evidence covers two math benchmarks with 150 questions each and models only up to 7B parameters, so it may not transfer to larger models or coding tasks. At 7B, Self-Refine and forced Reflexion remained **3.6–10.1 points below** baseline.",
  "source": {
    "name": "arXiv",
    "url": "https://arxiv.org/abs/2607.28576v1",
    "published_at": "2026-07-30T17:38:23.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Paper",
  "layer": "benchmark",
  "domains": [],
  "topics": [
    "agent-evals",
    "benchmark-integrity",
    "agent-reliability"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "The evidence covers two math benchmarks with 150 questions each and models only up to 7B parameters, so it may not transfer to larger models or coding tasks. At 7B, Self-Refine and forced Reflexion remained **3.6–10.1 points below** baseline."
  ],
  "connected_context": {
    "meaning": "This strengthens the case that gains attributed to reflection may instead come from spending more inference tokens or simply trying again. It turns equal-token repeated sampling into the minimum baseline for self-correction claims, while sharply limiting the conclusion to small open models and short math tasks. Long-horizon, larger-model, and coding-agent workflows still require separate tests.",
    "corpus_size": 297,
    "generated_at": "2026-07-31T10:07:56.623Z",
    "connections": [
      {
        "title": "Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2607.12962v1",
        "feed7_url": "https://feed7.dev/p/2607-12962v1-0q4i26c",
        "reason": "The placebo-controlled code-model study independently supports the same concern: retry scaffolding can match purported feedback-driven repair, so improvement alone does not show that reflection content caused it."
      },
      {
        "title": "Demystifying evals for AI agents",
        "source_name": "Anthropic",
        "source_url": "https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents",
        "feed7_url": "https://feed7.dev/p/demystifying-evals-for-ai-agents-1kh2tdz",
        "reason": "Anthropic’s distinction between pass@k and pass^k provides the evaluation vocabulary needed to report repeated independent attempts separately from reliability across attempts when comparing reflection workflows."
      },
      {
        "title": "Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=cO8qC6HBuBg",
        "feed7_url": "https://feed7.dev/p/vending-bench-long-horizon-agent-evals-lukas-petersson-andon-labs-0fu78nz",
        "reason": "Vending-Bench’s long-horizon drift and incentive effects mark an important boundary on transfer: equal-cost results from short math questions do not resolve whether reflection helps agents manage persistent state over extended tasks."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-07-30T17:38:23.000Z",
  "modified_at": "2026-07-30T17:38:23.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/2607-28576v1-09h2m1u",
    "json": "https://feed7.dev/p/2607-28576v1-09h2m1u.json",
    "markdown": "https://feed7.dev/p/2607-28576v1-09h2m1u.md"
  }
}