skip to content
The Weighted Average

Wire

ABSeeker takes a 4B search agent to 55.3%

ABSeeker lifted a Qwen3.5-4B search agent to 55.3% on BrowseComp with context management, up from 37.3% without it, after training on only 8,500 examples. The August 5 paper backtracks from known answers to score useful, redundant, and erroneous search steps separately; its 4B system then matched agents the authors describe as roughly 30B parameters. Teams building a discovery layer for agent tools and resources should test trajectory-level credit and context compaction before paying for a larger model, while treating the self-reported benchmark as a lead rather than deployment proof.