Sign InOpen Brain
arXivPaperNeeds Review

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

Skill-SP turns agent skills into units for verifiable self-play: generate tasks, solve them, then update the skill library from execution feedback. The abstract provides no per-benchmark effect sizes.

arXiv · Jul 24, 2026
Open Source Open MarkdownOpen JSON
Source Summary

**Skill Self-Play** combines a proposer, solver, and dynamic skill controller in a reinforcement-learning loop. Skills constrain each task to a verifiable scenario while dynamic routing expands task variety.

Practical Implication

For agent builders, the reusable pattern is to treat skills as both execution scaffolds and curriculum units: sample a skill, generate a harder task, evaluate execution, then revise the library from observed failures.

Agent-Ready Context
**Skill Self-Play** combines a proposer, solver, and dynamic skill controller in a reinforcement-learning loop. Skills constrain each task to a verifiable scenario while dynamic routing expands task variety.

For agent builders, the reusable pattern is to treat skills as both execution scaffolds and curriculum units: sample a skill, generate a harder task, evaluate execution, then revise the library from observed failures.

The paper reports gains on **tool-use and reasoning benchmarks**, but the supplied abstract gives no per-benchmark numbers, training costs, or evidence that the loop transfers cleanly to coding-agent workloads.
Context Map
agent#skills#tool-use#reasoning
Uncertainty
The paper reports gains on **tool-use and reasoning benchmarks**, but the supplied abstract gives no per-benchmark numbers, training costs, or evidence that the loop transfers cleanly to coding-agent workloads.