Wire
Latent rewards cut diffusion alignment compute 33x
Latent Reward Registers cut diffusion-alignment GPU hours by up to 33x while outperforming online reinforcement-learning baselines, according to the August 4 preprint. The method reads estimated terminal preference from intermediate noisy latents, then uses those dense gradients for on-policy distillation instead of paying for full rollout-heavy policy updates. For builders exploring the parallel-generation economics behind Google’s open diffusion-language model, the result makes intermediate reward prediction worth testing before adding more sampling compute.