Wire
ParamBench lifts exact tool-call match to 59.6%
ParamBench’s probe-guided framework raised average exact match on tool-call parameters from 19.7% to 59.6% across five open models and six external benchmarks. The paper builds its five-level benchmark from real cloud-network APIs, then uses a linear probe over hidden states to filter self-generated training calls or rerank candidates at inference. For teams extending Supabase’s outcome-based agent regression loop, the result is a useful warning: test exact argument values and cross-parameter dependencies, not merely whether an agent selected the right tool.