skip to content
The Weighted Average

Wire

Video research agent reaches 64% on 200 tasks

Video-DeepResearch-35B-A3B reached 64.0% average accuracy on a new 200-question, multi-hop video benchmark, versus 59.0% for Claude 4.5 Sonnet and 52.5% for GPT-5 under the authors’ protocol. The August 4 paper attributes the result to a staged tool policy that requires cross-frame visual grounding before web retrieval, countering agents’ tendency to skip video inspection and answer from text search or memorized knowledge. Builders extending Octobench’s finding that a harness changed five coding outcomes should treat tool order as an evaluated policy rather than prompt decoration, while regarding this small, newly introduced benchmark as directional until independently reproduced.