Wire
SVR cuts math reasoning to 2.99 turns
Self-Verifying Refinement reached 56.3% macro-average accuracy in 2.99 turns across seven math benchmarks with Qwen3.5-2B, beating the paper’s fixed ten-turn comparison while spending less inference time. The SVR paper trains the model to emit a correctness verdict and confidence score, then stop only when both clear its threshold rather than assigning every problem the same budget. For operators weighing smaller models against expensive reasoning contexts, the useful pattern is not “reason longer” but “teach the model when another pass is worth paying for.”