Agentic Engineering
Your Agent's Skill Library Gets Worse After 20 Entries
A 8,135-run study finds skill retrieval precision collapses from 29.6% to 3.3% as libraries grow from 5 to 100 — while task success barely moves.
Adding skills to a coding agent stops helping long before your library looks impressive. A controlled study across 8,135 trial records finds that as a candidate pool grows from 5 to 100 skills, the agent’s actual-use precision — the share of skills it invokes that are the right ones — falls from 29.6% to 3.3%, a ninefold collapse, while downstream task success moves only from 36.4% to 39.3%. The finding comes from “Demystifying Agent Skills: Why They Work—Until They Don’t”, a Princeton and UC San Diego collaboration evaluating Codex and Gemini CLI agents on Terminal-Bench and SkillsBench.
That gap between collapsing precision and flat success is the most useful thing in the paper, because it kills the intuitive mental model. Teams treat a skill library like a knowledge base whose value grows with entries. The data says it behaves more like a search index whose signal degrades with entries, while the agent muddles through anyway by reading several candidates and improvising. Recall stays high — 54.3% to 73.6% at a pool of 100, per the full text of the retrieval experiments — so the agent usually sees the right skill. It simply does not restrict itself to it. Coverage of the paper framed the retrieval collapse as the headline finding, with The Decoder noting that agents have a harder time finding the right instructions as libraries grow.
Skills are playbooks, not knowledge
The mechanism analysis explains why skills help at all. Across paired trajectories, procedural anchoring accounts for 65.7% of the cases where a skill improved a run, against just 4.5% for explicit knowledge injection. Skills stabilize what the agent does — setup sequences, tool ordering, output formats, verification steps — rather than telling it facts it lacked. The effect shows most clearly in execution-layer failures: environment and infrastructure failures fall from 5.3% of raw runs to 0.2% with skills, and output-format mismatches drop from 7.4% to 3.2%.
Representation is doing real work here, not just exposure to prior runs. When the same trajectories are handed to the agent as raw workflow memory instead of a distilled SKILL.md, performance is 6.06 percentage points worse in matched comparisons, with a 95% bootstrap interval of +0.76 to +11.36. Workflow memory drags along failed branches and verbose exploration; its distinctive penalty is timeout exhaustion, appearing in 10.6% of workflow runs versus 4.4% with skills. Distillation is the value-add, which is a direct empirical argument for the discipline of writing procedures down cleanly rather than dumping transcripts into context — the practice this paper described when Anthropic published its public skills catalog of reusable capability folders on the anthropics/skills repository.
Skills also create a failure surface that did not previously exist. The mode the authors call misapplied-or-ignored guidance appears in 10% of skill-arm runs, against 0.8% for raw execution — cases where the playbook is plausible, the agent follows it mechanically, and a condition that no longer holds carries over anyway. Skills do not repair bad algorithms either: logic errors persist at 7.4% with skills versus 8.3% raw, essentially unchanged.
What to change in your agent setup this week
The operational reading is about library hygiene, and it has a specific number attached. Semantic confusability, not raw size, is the dominant stressor for identification: embedding retrieval precision on similar pools falls from 70.5% at five candidates to 53.4% at one hundred, while dissimilar pools barely move, from 96.6% to 93.2%. Near-duplicate skills are the poison. Two entries named “run the test suite” and “verify tests before commit” cost more than they add. Library size is also a token-budget question, not only a retrieval one, and the cost of context is now moving with hardware — the subject of today’s lead on Nvidia’s 15% server increase.
Four moves follow. Prune aggressively toward one skill per distinct procedure and merge near-neighbors, because your retrieval problem scales with similarity rather than count. Scope libraries per repository or per task family instead of maintaining one global pool, which keeps the effective candidate set in the range where precision has not collapsed. Write skills as procedures — setup, tool order, checks, known failure modes — since that is where 65.7% of the measured benefit lives, and skip the background explainers. And instrument which skills your agents actually invoke, because the study’s core lesson is that availability and use are different measurements; a library nobody’s agent reads is documentation, not capability.
What could break this conclusion. The retrieval experiments use SkillsBench because it carries ground-truth task-to-skill annotations, and precision is scored against those labels — so a related, genuinely useful skill counts as a false positive. The authors say so plainly: exact ground-truth invocation is “neither sufficient nor necessary” for success, which is why success stayed flat while precision fell. Read the 3.3% as a measure of targeting discipline, not of wasted effort. The RQ4 experiments also swap in GPT-5.4 for the Codex pairing because the original model was unavailable, so those absolute values are within-pairing comparisons rather than cross-experiment ones.
The verdict: skills remain worth building, and the distillation step is where the gain is — a conclusion consistent with earlier evidence that agent-team structure matters less than shared artifacts, as the 1,902-run study on multi-agent coordination tokens found when designated coordinators changed nothing. But treat the library as a retrieval system with a maintenance cost, not an archive that compounds for free. The evidence that would change the verdict is a benchmark where success rates, not just precision, degrade with pool size; this study explicitly did not find one, and until someone does, the cost of a bloated library is targeting noise rather than outright failure.
Sources
- arXiv — Demystifying Agent Skills: Why They Work—Until They Don’t
- arXiv — full text of the agent skills study, including retrieval and mechanism results
- GitHub — anthropics/skills, the public catalog of reusable capability folders
- The Decoder — study explains why AI agents benefit from skills and when they fail