skip to content
The Weighted Average

Wire

AgentOPSD reaches 89.1% without extra rollouts

AgentOPSD reached 89.1% success on ALFWorld with Qwen2.5-7B while adding neither a critic nor extra rollouts. The new preprint converts teacher-student log-probability gaps into turn-level evidence, recursively updates a Bayesian belief across an episode, and reports gains over GRPO and self-distillation baselines on ALFWorld, WebShop, and Search-QA; its linked repository says the full training code is still forthcoming. For teams extending the evidence that harness design changes agent outcomes, the method is worth filing as a potentially cheaper way to credit pivotal actions, but the score should remain provisional until the code enables reproduction.