Wire
Skill-entropy RL doubles Qwen3-4B's score
Skill-Entropy RL raised Qwen3-4B-Instruct’s Skill2-Bench score from 34.4% to 68.4% by rewarding the model for choosing the right reasoning skill at each step, not only for reaching the right answer. The research paper builds the benchmark from 558 skills across nine domains and finds that all 12 evaluated models lose accuracy when a familiar skill sits inside a cross-skill task; the authors also released the training code. For teams evaluating multi-stage agents, that makes skill transitions a separate failure surface from the model-and-harness effects measured in Octobench’s coding-agent comparison.