skip to content
The Weighted Average

Wire

Explicit skills tie context on 500 agent tasks

ContinualSkillBench tested agent learning across five domains with 100 interconnected subtasks each, and found explicit skill maintenance performed only comparably to in-context learning on average. The researchers report that reusable skills help selectively, especially when tasks demand repeatable procedures or exact outputs, while weaker models accumulate larger, more fragmented task-specific libraries. Builders extending the harness lessons in Octobench’s five-task performance swing should benchmark retained skills against an equal-context baseline before paying the complexity and token cost of a permanent skill store.