Wire
GRAFT scores step-level credit in agent training
GRAFT, a graph-based reinforcement-learning method, reports 25.5% and 16.7% higher WebShop average success than GRPO for 1.5B- and 7B-parameter models, with results averaged over three random seeds. In the GRAFT paper, the method merges rollout trajectories into a graph to assign step-level credit, and reports gains over GiGPO and GraphGPO on both ALFWorld and WebShop. That makes step-level credit worth testing before adding more agents; harness studies already show topology can move token cost without changing model weights.