Wire
Only three small-model tests clear a 20% risk budget
A study of 25,168 local predictions from 11 instruction-tuned models found that only three of 22 model-task pairs earned certified autonomy at a 20% risk budget, and none did at 10%. The certified-deferral paper used a 200-question calibration set and found that even Platt scaling as low as 0.02 expected calibration error did not make most systems safe to run without escalation. Operators following the case for routing security work to a small model should therefore treat calibrated confidence as a signal, not a release gate: certify the risk-coverage policy on the deployed task and defer the rest.