skip to content
The Weighted Average

Wire

HIVE finds voice prompts lose 9.7 points

HIVE’s 550,000 scored generations found voice-transcription perturbations cut instruction-tuned model accuracy by an average 9.7 points, versus 3.0 for keyboard errors, according to the August 4 preprint. Across five models and six benchmarks, more reasoning budget almost entirely repaired keyboard corruption but not spoken-register damage; the worst rewrite, compressed speech, lost 24.1 points. Teams adding voice to coding agents should treat transcription as a tested system component—not a transparent pipe—alongside the cost and adoption case for hands-free Codex.