skip to content
The Weighted Average

Wire

Calibrated 30B model hits 82.6% alert accuracy

A fine-tuned 30B reasoning model reached 82.6% test accuracy on human-labeled Windows endpoint alerts, improving high-confidence benign recall 43.0% and malicious recall 18.3% over a direct-label classifier. The authors’ preprint says chain-of-thought weakened the label-token probabilities automation normally trusts, so they trained a separate calibrator on the full reasoning trace; an untrained confidence judge drove high-confidence recall to zero. Security teams following Microsoft’s specialized-model routing pattern should budget for calibrated abstention as a distinct component—not mistake a model’s verbal confidence for a safe automation threshold.