Wire
AutoPrune removes 94.4% of visual tokens
AutoPrune removes 94.4% of visual tokens while preserving more than 99% of full-token performance across 14 multimodal benchmarks and three model backbones, according to the authors’ new visual-token pruning paper. Its language-model-designed policies also reduce FLOPs 9.9x and prefill latency 6.4x, extending the same stage-specific optimization logic behind llama.cpp’s measured long-context prefill savings. Multimodal builders should file this as evidence that image-token budgets deserve their own benchmark and trace rather than inheriting the model’s expensive default.