Sign InOpen Brain
arXivPaperNeeds Review

Boosting LLM Exploration via Weak-Model Guidance in RLVR

Feeding a model partial reasoning from a weaker model during RLVR preserved exploration and improved large-k coverage. This suggests diversity can come from cross-model perturbation, not only regularization.

arXiv
Open Source Open MarkdownOpen JSON
Source Summary

The method gives an RLVR target model partial reasoning prefixes from a **smaller, weaker model**. These unfamiliar prefixes interrupt overconfidence, preserve generative diversity, and outperform vanilla RLVR across multiple math benchmarks.

Practical Implication

Model trainers should consider cross-model prefixes when RLVR collapses entropy or narrows pass@k coverage. The intervention requires **no additional SFT**, custom reward design, or elaborate prompting.

Agent-Ready Context
The method gives an RLVR target model partial reasoning prefixes from a **smaller, weaker model**. These unfamiliar prefixes interrupt overconfidence, preserve generative diversity, and outperform vanilla RLVR across multiple math benchmarks.

Model trainers should consider cross-model prefixes when RLVR collapses entropy or narrows pass@k coverage. The intervention requires **no additional SFT**, custom reward design, or elaborate prompting.

Reported gains become larger as **k scales up**, but the supplied abstract contains no effect sizes, model names, or compute costs. Evidence is limited to mathematical benchmarks, so transfer to coding-agent reasoning remains open.
Context Map
modelresearch#reasoning
Uncertainty
Reported gains become larger as **k scales up**, but the supplied abstract contains no effect sizes, model names, or compute costs. Evidence is limited to mathematical benchmarks, so transfer to coding-agent reasoning remains open.