Wire
OmniPack retains 98% at one-sixth the FLOPs
OmniPack preserved 98.0% of Qwen2.5-Omni-7B’s original performance while reducing inference to 16.7% of the original FLOPs across the authors’ evaluation, and retained 92.9% at a 6.8% compute budget. The August 4 preprint describes a training-free pipeline that compresses redundant audio and visual tokens before the language model, then refines them using the query after multimodal interaction; experiments cover five benchmarks and three omni-model backbones. For teams weighing the deployment footprint behind Inkling-Small’s open-weight control premium, the result makes token compression worth an equal-quality pilot before paying for more inference hardware, with long-context recall and rare-event retention as the break tests.