Sign InOpen Brain
arXivPaperNeeds Review

Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs

Small proxy models may be enough to choose an SFT-versus-RL annotation split: the paper finds broad near-optimal ranges that transfer to larger models.

arXiv
Open Source Open MarkdownOpen JSON
Source Summary

The study replaces a single best SFT–RL split with a range of allocations near peak performance. These regions remained wide at **2–10% tolerance**, widened with model scale, and transferred from small proxies to larger targets.

Practical Implication

If you post-train a model for agent workloads, search the allocation space on a small proxy first, then carry the near-optimal range into the expensive run. Budget using annotation cost, because unequal **SFT and RL data costs** shift that range.

Agent-Ready Context
The study replaces a single best SFT–RL split with a range of allocations near peak performance. These regions remained wide at **2–10% tolerance**, widened with model scale, and transferred from small proxies to larger targets.

If you post-train a model for agent workloads, search the allocation space on a small proxy first, then carry the near-optimal range into the expensive run. Budget using annotation cost, because unequal **SFT and RL data costs** shift that range.

The material reports consistency across tasks, model families, and **off-policy and on-policy RL**, but gives no concrete savings or universal allocation ratio. Transfer still needs validation for a builder’s own data and evaluation target.
Context Map
modeldata#model-selection
Uncertainty
The material reports consistency across tasks, model families, and **off-policy and on-policy RL**, but gives no concrete savings or universal allocation ratio. Transfer still needs validation for a builder’s own data and evaluation target.