Wire
QaDiT opens a 159M-parameter audio baseline
QuarkML released QaDiT, an Apache-2.0 text-to-audio diffusion transformer with roughly 159 million parameters that generates fixed 10.24-second, 16 kHz mono clips from a caption. The QaDiT model card says the checkpoint trained for 24,000 steps on 45,000 captioned clips and exposes raw latents and denoising predictions, but warns that its timbre is coarse, complex-prompt adherence is weak, and it is unsuitable for speech or music production. Like the control tradeoff in Inkling-Small’s open-weight release, QaDiT is most useful as an inspectable research baseline: audio builders should benchmark it as a sampler testbed, not mistake downloadable weights for production quality.