arXivPaperNeeds Review
Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs
Small proxy models may be enough to choose an SFT-versus-RL annotation split: the paper finds broad near-optimal ranges that transfer to larger models.
arXiv
Source Summary
The study replaces a single best SFT–RL split with a range of allocations near peak performance. These regions remained wide at **2–10% tolerance**, widened with model scale, and transferred from small proxies to larger targets.
Practical Implication
If you post-train a model for agent workloads, search the allocation space on a small proxy first, then carry the near-optimal range into the expensive run. Budget using annotation cost, because unequal **SFT and RL data costs** shift that range.
Agent-Ready Context
The study replaces a single best SFT–RL split with a range of allocations near peak performance. These regions remained wide at **2–10% tolerance**, widened with model scale, and transferred from small proxies to larger targets. If you post-train a model for agent workloads, search the allocation space on a small proxy first, then carry the near-optimal range into the expensive run. Budget using annotation cost, because unequal **SFT and RL data costs** shift that range. The material reports consistency across tasks, model families, and **off-policy and on-policy RL**, but gives no concrete savings or universal allocation ratio. Transfer still needs validation for a builder’s own data and evaluation target.
Context Map
modeldata#model-selectionUncertainty
The material reports consistency across tasks, model families, and **off-policy and on-policy RL**, but gives no concrete savings or universal allocation ratio. Transfer still needs validation for a builder’s own data and evaluation target.