skip to content
The Weighted Average

Wire

Beacon lifts visual reasoning 5.9 points

Beacon averaged 58.98% across 13 visual benchmarks, 5.89 percentage points above its Qwen3-VL-8B base and 2.31 points above the strongest agentic baseline. The authors’ preprint also found that tools raised Beacon’s five-benchmark average from 51.57% to 53.53%, while its training reward taught the model to distinguish questions that warranted tool calls from those it could answer directly. Teams building multimodal agent workflows around frontier models should measure tool harm as well as tool gain: Beacon’s 9.29-point gain on rescued answers still came with a 6.15-point loss on answers the text-only path had already solved.