Wire
VAD lifts visual distillation 3.05 points
Visual Attribution Distillation lifted a 9B model’s six-benchmark average from 76.88% to 79.93%, a 3.05-point gain over direct privileged-view distillation. The VAD preprint attributes teacher corrections by comparing evidence-present and evidence-degraded views, and reports 96.3 GPU-hours—11.1% more step time than the direct baseline—using the same 6,241 synthetic examples; the training and evaluation code is public. Teams weighing the control premium of open multimodal models should file the method as a reproducible way to separate visual evidence from teacher priors, while remembering that its published gains remain preprint results judged through a GPT-OSS-120B evaluation pipeline.