Wire
Reasoning study releases 22M inference traces
A test-time scaling study assembled more than 22 million full reasoning traces while dividing inference into three regimes: one evolving trajectory, completed-candidate aggregation, and search over unfinished prefixes, according to the August 4 preprint. Its practical warning is that an accuracy score and nominal token budget do not identify the system tested; prompts, decoders, verifiers, stopping rules, uncertainty, and compute accounting all belong in the result. Model evaluators should apply that protocol-level ledger alongside the evidence that a harness alone can swing five coding tasks before buying a leaderboard claim.