Wire
DEFT-RLVR lifts driving decisions to 77.9%
DEFT-RLVR raised Qwen3-VL-8B trajectory-selection accuracy from 28.1% to 77.9% on AD-MCQ-500 by making the model reason about a driving scene before revealing candidate paths, according to the August 4 release of its code, checkpoints, and dataset. Multiple-choice planning is not closed-loop driving safety, but the result gives teams extending simulation-heavy robot evaluation funnels a reproducible way to test whether a vision-language model derived its decision from scene evidence rather than rationalized an exposed future trajectory.