Wire
DiffusionGemma retrains on under 10% of the tokens
Google DeepMind converted Gemma 4 into DiffusionGemma using fewer than 10% of the original model’s training-token budget rather than pretraining a diffusion model from scratch. The technical report newly surfaced in coverage says the 25.2-billion-parameter model averages about 20 generated tokens per forward pass and roughly 1,500 output tokens per second on one H100, though its single-user speed advantage narrows as concurrency rises. For teams following DiffusionGemma’s parallel-decoding economics, the reusable fact is the conversion budget: architecture experiments may be financeable as post-training projects instead of full foundation-model runs.