Wire
HySparse2 cuts 1M-token prefill 5×
HySparse2 cuts 1M-token prefill FLOPs by 5.02× against hybrid sliding-window attention, while its KV cache measures 2.69 GB versus 12.09 GB in the comparison reported by the arXiv paper on its two-level KV sharing. On an 80B-A3B mixture-of-experts model, the authors also report a 19.81-point RULER-v2 lift over HySparse after post-training. The paper’s agent-history cost lesson still applies: treat this as research evidence for long-horizon inference, not a production benchmark for your serving stack.