skip to content
The Weighted Average

Wire

Taste-Bench puts agent taste below 60%

A Microsoft-affiliated research team introduced Taste-Bench, a 502-question test of long-horizon agent decisions, and the best evaluated model reached only 59.7% accuracy. The benchmark hides what happens after each decision fork, so a high end-to-end score can still mask poor early judgment; more reasoning budget did not improve accuracy. Builders extending the archive’s OSWorld completion baseline should log intermediate choices, not only final task success.