skip to content
The Weighted Average

Wire

Clinical audit misses all 48 proxy failures

KAISEN’s clinical-model audit correctly diagnosed 144 of 144 controlled cases but recovered zero of 48 model-driven cases when a proxy was misspecified—and gave no signal that the diagnostic had failed. The five-phase stress test also found per-group thresholds reduced equalized-odds disparity in all 48 held-out runs, while cohort seeds drove 27 of 27 false drift alarms; every result is synthetic and does not establish clinical validity. Builders operating near ChatGPT Health’s external-validation gap should test audit pipelines against deliberately known failure mechanisms instead of treating a clean dashboard as evidence that subgroup risk is controlled.