skip to content
The Weighted Average

Wire

Cursor lifts MoE training throughput 41%

Cursor open-sourced Mixture-of-Kittens after the MoE megakernel raised end-to-end tokens per second 1.41x in its production training stack across multiple NVL72 racks. Cursor says the kernel fuses MoE communication and computation and now powers Composer training across tens of thousands of GPUs, though the repository requires NVIDIA Blackwell hardware and workload-specific tuning. For teams studying the nine-H200 floor behind Qwen3.8-Max’s giant open-weight preview, the release is a useful reminder that rack-level communication—not just parameter arithmetic—can dominate the cost of training sparse models.