arXivPaperNeeds Review
Boosting LLM Exploration via Weak-Model Guidance in RLVR
Feeding a model partial reasoning from a weaker model during RLVR preserved exploration and improved large-k coverage. This suggests diversity can come from cross-model perturbation, not only regularization.
arXiv
Source Summary
The method gives an RLVR target model partial reasoning prefixes from a **smaller, weaker model**. These unfamiliar prefixes interrupt overconfidence, preserve generative diversity, and outperform vanilla RLVR across multiple math benchmarks.
Practical Implication
Model trainers should consider cross-model prefixes when RLVR collapses entropy or narrows pass@k coverage. The intervention requires **no additional SFT**, custom reward design, or elaborate prompting.
Agent-Ready Context
The method gives an RLVR target model partial reasoning prefixes from a **smaller, weaker model**. These unfamiliar prefixes interrupt overconfidence, preserve generative diversity, and outperform vanilla RLVR across multiple math benchmarks. Model trainers should consider cross-model prefixes when RLVR collapses entropy or narrows pass@k coverage. The intervention requires **no additional SFT**, custom reward design, or elaborate prompting. Reported gains become larger as **k scales up**, but the supplied abstract contains no effect sizes, model names, or compute costs. Evidence is limited to mathematical benchmarks, so transfer to coding-agent reasoning remains open.
Context Map
modelresearch#reasoningUncertainty
Reported gains become larger as **k scales up**, but the supplied abstract contains no effect sizes, model names, or compute costs. Evidence is limited to mathematical benchmarks, so transfer to coding-agent reasoning remains open.