skip to content
The Weighted Average

Wire

Reflection wins none of 36 equal-cost tests

Self-reflection and related reasoning schemes reliably beat repeated sampling in none of 36 equal-token-cost comparisons across seven methods, three open-model sizes, and two math benchmarks. The researchers’ paired evaluation found 10 comparisons reliably worse and all 18 self-inspection comparisons negative; Self-Refine and forced Reflexion remained 3.6 to 10.1 points behind the baseline at 7B parameters. For operators confronting the end of unconstrained token spending, the result argues for measuring every critique and rewrite token, then requiring reflection loops to beat plain repeated sampling at the same all-in budget.