skip to content
The Weighted Average

Wire

Chimera cuts video-model compute 7.3x

Chimera’s complete visual-diffusion system was 7.3 times as compute-efficient by pretraining loss as a matched full-attention Wan-2.1 2B baseline. The 40-page preprint combines linear Kimi Delta Attention, periodic global attention, local convolutions, and sparse experts; its 11-billion-parameter model activates 2 billion parameters and extends 5-second training clips to 30-second videos with 6.5% FID degradation in the final five seconds. For video teams tracking the gap between generation claims and production economics, hybrid attention now merits a controlled benchmark before another full-attention training run—though loss efficiency still needs translation into output quality and wall-clock cost.