# Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B

Source: [arXiv](https://arxiv.org/abs/2607.28576v1)  
Feed7 permalink: https://feed7.dev/p/2607-28576v1-09h2m1u  
Published: 2026-07-30T17:38:23.000Z  
Trust: Needs Review (needs_review)

## Why Included

At equal generated-token cost, repeated sampling matched or beat seven reflection, critique, selection, and debate methods. Agent evals should budget every generated token, not compare against one-shot baselines.

## Source Summary

Researchers tested **seven methods** on 1.5B, 3B, and 7B open models across two math benchmarks, counting critique, reflection, debate, and checking tokens. Across **36 paired comparisons**, none reliably beat repeated sampling at equal cost.

## Practical Implication

For agent workflows, benchmark reflection loops against repeated independent attempts using the same measured token budget. Self-inspection deserves particular scrutiny: **all 18 comparisons were negative**, and ten results were reliably worse.

## Agent-Ready Context

Researchers tested **seven methods** on 1.5B, 3B, and 7B open models across two math benchmarks, counting critique, reflection, debate, and checking tokens. Across **36 paired comparisons**, none reliably beat repeated sampling at equal cost.

For agent workflows, benchmark reflection loops against repeated independent attempts using the same measured token budget. Self-inspection deserves particular scrutiny: **all 18 comparisons were negative**, and ten results were reliably worse.

The evidence covers two math benchmarks with 150 questions each and models only up to 7B parameters, so it may not transfer to larger models or coding tasks. At 7B, Self-Refine and forced Reflexion remained **3.6–10.1 points below** baseline.

## Connected Context

Feed7 judgment across 297 accumulated Signals:

This strengthens the case that gains attributed to reflection may instead come from spending more inference tokens or simply trying again. It turns equal-token repeated sampling into the minimum baseline for self-correction claims, while sharply limiting the conclusion to small open models and short math tasks. Long-horizon, larger-model, and coding-agent workflows still require separate tests.

- [Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models](https://feed7.dev/p/2607-12962v1-0q4i26c) — The placebo-controlled code-model study independently supports the same concern: retry scaffolding can match purported feedback-driven repair, so improvement alone does not show that reflection content caused it.
- [Demystifying evals for AI agents](https://feed7.dev/p/demystifying-evals-for-ai-agents-1kh2tdz) — Anthropic’s distinction between pass@k and pass^k provides the evaluation vocabulary needed to report repeated independent attempts separately from reliability across attempts when comparing reflection workflows.
- [Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs](https://feed7.dev/p/vending-bench-long-horizon-agent-evals-lukas-petersson-andon-labs-0fu78nz) — Vending-Bench’s long-horizon drift and incentive effects mark an important boundary on transfer: equal-cost results from short math questions do not resolve whether reflection helps agents manage persistent state over extended tasks.

## Context Map

- Layer: benchmark
- Domains: None
- Topics: agent-evals, benchmark-integrity, agent-reliability

## Uncertainty

- The evidence covers two math benchmarks with 150 questions each and models only up to 7B parameters, so it may not transfer to larger models or coding tasks. At 7B, Self-Refine and forced Reflexion remained **3.6–10.1 points below** baseline.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
