Wire
CoBa cuts reasoning tokens 58.9%
CoBa matches best-of-16 majority voting within 0.01 percentage points while using 58.9% fewer parameter-weighted tokens across 3,129 math and symbolic-reasoning evaluations, according to the compute-balanced routing paper. The policy spends cheap verification broadly, reserves stronger verification for uncertain candidates, and reaches 85.13% macro accuracy—evidence for the cost-per-resolved-task discipline already replacing flat token prices. Reasoning-system operators should benchmark generation, verification, and stopping as separate budget choices instead of buying a larger sample count by default.