Wire
Cue cuts voice-agent polish latency 44%
Cue moved its voice agent’s text-polish step to Gemma 4 E4B on-device and cut median latency 44%, from 876 to 488 milliseconds, across 227 real-voice samples. Dictation per active beta user rose about 30% after the switch, while the step’s marginal inference cost fell to zero; the result extends the local-compute case behind llama.cpp’s 4.92-second DeepSeek V4 prefill gain. Builders should benchmark narrow, frequent transformations locally before renting frontier inference, because the smaller model can win simultaneously on latency, privacy, and unit cost.