skip to content
The Weighted Average

Wire

Programmatic tools beat JSON in 11 of 14 models

Programmatic tool calling matched or beat native JSON calls in 11 of 14 language models tested on BFCL v4, according to a new comparative study that exposed tools as typed Python stubs and executed each script within one agent turn. The authors report a 10.6% improvement for the GPT-5.6 family, parity or better in 13 of 14 models during parallel fan-out, and stable results under context rot while JSON performance fell 2.3% on average. For builders, the finding sharpens Octobench’s evidence that the harness changes agent outcomes: test the tool-call representation alongside the model, especially when workflows chain or parallelize many functions.