skip to content
The Weighted Average

Wire

TutorMoments tests 7,280 AI-tutor attempts

AI2’s TutorMoments benchmark tests 7,280 AI-tutor attempts across 520 teacher-identified decision points, seven models, and two prompts, using 462 de-identified math-tutoring transcripts. The released dataset and benchmark card says minimally prompted models tend to over-help and rarely press students to reason more deeply; explicit scaffolding-versus-rigor instructions improve every model, but the synthetic students mean the scores measure tutor behavior rather than learning. For education-product teams confronting the institutional-readiness gap in the 2026 AI Index, the practical lesson is to evaluate when an assistant withholds help—not merely whether its answers are correct.