skip to content
The Weighted Average

Wire

Fisher-R1 lifts statistical agent reliability 21%

Fisher-R1-14B improved single-trial success by 21% relative over DeepSeek-V4-Pro on P-Bench, a 425-task benchmark requiring agents to choose a statistical method, compute a p-value, and draw a conclusion. The preprint reports gains up to 26% on the hardest tasks and says the open-weight agent outperformed GPT-5.4 and other tested baselines after reinforcement learning on verified statistical rewards. For teams using scientific agents in regulated workflows, the signal is not “trust the model”: make assumptions, p-values, and conclusion validity explicit review artifacts rather than accepting code execution as evidence of sound inference.