skip to content
The Weighted Average

Wire

GPT-Live makes its old p50 the new p95

OpenAI says GPT-Live’s new Go media stack made frame delivery smooth enough that its p95 matched the previous Python asyncio system’s p50, while Instant Connect can start a session with one UDP packet. The company’s production engineering account describes a full-duplex voice model that keeps media on a dedicated path and sends reasoning and tool calls across an asynchronous boundary. For builders following OpenAI’s move from voice commands toward hands-free agents, the useful architecture lesson is that tail latency, session state, and tool isolation now matter as much as model response speed.