skip to content
The Weighted Average

Wire

Skill retrieval misses one in four agent queries

A 690-skill agent library put the correct skill in its top five for 73.5% ± 8.0% of 117 realistic queries, according to a new hybrid-retrieval study. Replacing ranked results with knowledge-graph neighbors lowered accuracy by 11.2 points at the same token budget, while author-written queries overstated hit rate by as much as 44 points. For builders extending the evaluation discipline required when agents cross scope, the result argues for testing skill routing on user phrasing—not catalog language—and keeping a fallback for the roughly one-quarter miss rate.