Wire
ByteDance halves pacing errors with SeedRealtime
ByteDance says SeedRealtime, its native audio-visual full-duplex model, cut conversational pacing problems by 50% against cascaded systems in an end-to-end human evaluation. The SeedRealtime launch announcement says one architecture continuously fuses audio, video, and text, decides when to speak without an external voice-activity detector, and can proactively invoke tools; it is fully rolled out, although ByteDance disclosed neither the evaluation sample size nor API access. Builders following the shift toward full-duplex voice as an agent interface should file away the timing gain, but treat the model as a consumer-product signal until reproducible benchmarks, pricing, and developer access arrive.