skip to content
The Weighted Average

Wire

ReToken lifts visual retrieval 13.4 points

ReToken improved Qwen3VL-8B by 13.4 points on Visual Haystacks while keeping both training and long-video inference on one H100. The researchers’ preprint and open code describe one learned embedding that retrieves a sparse, query-relevant subset from a prefilled visual KV cache; the same method added 12.4 points to InternVL3.5 and transferred zero-shot to LVBench for an 8-point gain. For teams already weighing the cost of sparse open models, the result is a reason to test retrieval over cached visual tokens before paying to process every frame through the full context window.