Qwen research uses self play to improve AI tool skills

Self-play without a leash collapses into noise or a toy sandbox. Qwenโs application team argues the fix is a skill library that co-evolves with the model. In Skill Self-Play (Skill-SP), a proposer, solver, and skill controller train together so the curriculum stays hard enough to teach and clean enough to trust.
Stake: open-ended tasks with verifier-grade feedback, not synthetic prompts that look diverse and grade themselves wrong.
How proposer, solver, and skill library co-evolve
Skill-SP runs as a reinforcement loop. The proposer samples a skill package and writes a task plus a hidden verification contract. The solver only sees the prompt and must produce a correct, schema-valid answer. The skill controller watches which packages yield frontier difficulty, then refines, prunes, or induces new skills from an open-ended exploration stream.
A skill bundles routing metadata, generation rules, hints, examples, and executable validators. It injects structure before synthesis and checks contracts after. The proposer is rewarded for medium-difficulty, valid tasks (frontier around 50% solver success), not for inventing unsolvable traps. Policies update with GRPO across five iterations.
The dual stream matters. Half the solverโs pool comes from skill-conditioned generation; half from unconstrained exploration that can mint new patterns. Without the mix, the loop narrows into templates or floods itself with garbage.
Avoiding narrow environments and noisy synthesis
Environment-bound self-play (code executors, game sims, retrieval sandboxes) gives crisp rewards and tiny task spaces. Unguided self-generation widens coverage, then fails quality control: ill-posed items, template overfitting, synthetic collapse. Post-hoc filters catch format errors. They do not steer the generator.
Skills sit in the middle. Each package is deep and verifiable in one scenario; dynamic routing keeps variety without stuffing raw trajectories into an ever-fatter proposer prompt. Ablations on Qwen3-4B-Instruct back that claim: unguided self-play drops overall tool-call score by 2.6 points versus full Skill-SP; frozen skills and uniform routing also lag. Freeze the proposer or the feedback solver and the frontier signal goes stale.
API-Bank, BFCL, and ZebraLogic gains
Benchmarks cover tool calling (API-Bank L1โL3 and four BFCL slices) and constraint reasoning (ZebraLogic). Backbones span Qwen3-4B-Instruct, Qwen3-8B, Ministral-3-8B/14B, and Granite-4.1-3B. The same checkpoint seeds proposer and solver; no stronger teacher in the main setup.
Headline deltas: up to +42.9 absolute on tool use and +12.0 on logical reasoning. Competent Qwen and Granite bases pick up steadier 2.8โ6.5 point tool-call lifts. The shock is the rescue case. Ministral-3-8B was too misaligned to synthesize valid tasks alone, so unguided self-play stalled; Skill-SP still turned it around because skills supplied structure the base could not invent. On ZebraLogic, Ministral-3-14B gained up to 12 grid points overall. Unguided self-play could not bootstrap valid puzzles, so it was not scored in reasoning.
Diagnostics show the skill stream sits closer to the frontier (mean solver success ~0.57) than unguided pools, while the mixed set covers more embedding space. Across five rounds the library induces ~20 new tool packages per iteration.
Whatโs in the public skill-self-play repo
Code is public at Qwen-Applications/skill-self-play. Teams can inspect the skill schema, curriculum builder, and GRPO loop instead of reverse-engineering a blog demo.
Analysis: Skill-SP bets that training-time skills beat inference-time skill stuffing. If the library is the curriculum engine, labs stop choosing between sandbox safety and open-ended chaos. Weak bases still need a competence floor, and mix heuristics may need retuning. For tool-calling and puzzle-hard reasoning, co-evolving skills look like Qwenโs sharpest self-play recipe on arXiv this cycle.



