Wire
YODAS v3 opens 1.1M hours of speech
ESPnet released YODAS v3, a 1.1 million-hour open multilingual speech dataset spanning 100-plus languages, with 48 kHz audio, word- and utterance-level timestamps, and timestamped English translations for more than half of the multilingual data. The Hugging Face release says more than 70% of recordings have at least two distinct channels, while the dataset card lists a CC BY 3.0 license, so voice builders get scale but still need provenance and license review. Teams measuring voice-agent context costs should file YODAS v3 as a data-availability shift—not proof that a model trained on it will generalize across accents or noisy production calls.