Wire
AMD cuts KV cache 92% with Zebra-HyLo
AMD released 14 Zebra-HyLo checkpoints that the ROCm announcement says cut KV-cache memory by at least 92%. The post-training recipe extends usable context up to 32× and serves up to 2M tokens on eight MI300X GPUs, while the checkpoints use Apache-2.0; the 64K recipe still carries a non-commercial training-data caveat. For teams hitting long-context memory ceilings, the cost of retaining agent history is the right comparison: this is a serving-memory escape hatch, not proof that every workload should move to AMD.