skip to content
The Weighted Average

Enterprise AI & Work

OSWorld's 20.6% Agent Score Is Already Stale

Two Fable 5.1 completion reports differ by more than 3.3 points. OSWorld buyers need versioned results, not a recycled paper headline.

Blank paper, a pen, a keyboard, and a potted succulent on a white desk
Blank paper, a pen, a keyboard, and a potted succulent on a white desk. Photograph by Mediamodifier

Snorkel AI’s September 3 OSWorld 2.0 reading-group report says Claude Fable 5.1 exceeded 45% full task completion, while Anthropic’s release table reports 41.7% strict completion. That leaves more than 3.3 percentage points between published accounts of the same model—not evidence of a performance change, but a reason to demand versioned evaluation records before buying computer-use automation.

This backfill was reconstructed on September 7, 2026, from records available by September 4, 2026.

A benchmark can outlive its headline

The Snorkel reading-group account introduces OSWorld 2.0 through a sobering paper-era result: even the strongest system tested completed only about one task in five. Later in the same account, the reported frontier moves. Presenter Mengqi Yuan describes Fable 5.1 as exceeding 45% binary completion and 60% partial credit. Repeating the introductory statistic as September’s best available result would erase the most important qualification in the source.

The distinction is between a paper and a living benchmark, not between a true number and a false one. The arXiv paper records 20.6% binary completion and 54.8% partial score for Opus 4.8 with maximum thinking and batched tool calls at a 500-step limit. Its listed second version dates to July 13. The result describes that evaluated system under those conditions; it is not a perpetual ceiling on every later model using the benchmark’s name.

There is a second discrepancy worth quantifying. Anthropic’s Fable 5.1 launch table reports 41.7% strict completion, versus the reading group’s greater-than-45% account. Subtract the two sourced figures: 45 − 41.7 = 3.3, so the published reports differ by more than 3.3 percentage points. This measures disagreement between accounts, not a model improvement or a verified benchmark error. Snorkel supplies no exact figure or complete matched settings for its newer claim, so the available evidence cannot reconcile them.

Anthropic explicitly warns that its scores use the authors’ August task release and are not directly comparable with earlier published OSWorld 2.0 results. It also says production-safeguard interventions score zero on this benchmark. Those disclosures explain why release and grading metadata matter; they do not establish the cause of the discrepancy with the talk. The vendor table is the precise first-party result to quote, while the talk’s lower-bound claim remains attributed and unresolved.

The benchmark remains relevant because of the work it asks agents to finish. The official project description specifies 108 long-horizon workflows, with a median human completion time around 1.6 hours. Its tasks require reconciling information across applications, recovering unstated context, and responding when the environment changes. These are not merely longer sequences of otherwise independent clicks. A late piece of information can invalidate a plan that looked correct when the agent began.

One representative reimbursement workflow requires the agent to follow guidance, assemble receipts and account evidence, recover missing identity information, notice an update, and submit the actual claim. The official failure examples describe preparation consuming the available horizon while the final submission remains unfinished. That distinction matters economically: collecting evidence can be useful assistance, but it is not equivalent to completing the business transaction.

Partial credit exists to describe that useful intermediate work. It should not be translated directly into the fraction of employee labor eliminated. Snorkel’s discussion explains that checkpoints are weighted and judged against the final machine state, rather than requiring every earlier checkpoint to be correct before a later one earns credit. A partially successful artifact and a safely completed workflow answer different operating questions. Buyers should keep both measures, with separate labels, rather than select whichever makes automation look more impressive.

Freeze the environment before you compare the agents

The first switching decision is methodological. An enterprise team evaluating unattended computer use should move from unversioned leaderboard screenshots to reproducible runs of its own critical workflows. That does not require rejecting the public benchmark. It means treating the benchmark as a diagnostic framework and demanding enough metadata to understand whether two reported results actually describe the same test.

The project’s published repository documentation makes versioning an explicit requirement. Its dated August 8 release updates code, task files, task assets, and mocked websites, and it warns against mixing releases or substituting development branches for supported versions. That is a concrete source of possible disagreement between nominally identical OSWorld scores. Record those components alongside the model, effort setting, tool access, and step budget before drawing a trend line.

This builds on our earlier argument for auditing benchmark questions. A score deserves attention only after the evaluator knows what its questions test and how the run was graded. For computer-use work, the environment is part of the question. If its state changes across model comparisons, a precise percentage can conceal a less precise experiment.

The monetary cost cannot be honestly reduced to a price per completed task from these sources. Snorkel reports that GPT models use fewer output tokens while scoring lower, but the talk does not provide a complete matched bill for the newer Fable result. The published reference budget allows 500 steps; that is an execution limit, not a dollar amount. A procurement team should collect actual model usage, elapsed time, and human recovery effort rather than infer a price from the completion percentage alone.

There is also a safety denominator. The reading-group account describes checks for sensitive-information leakage and harmful side effects. A workflow that reaches the desired endpoint while damaging unrelated files should not count as an operational success. Keep those outcomes separate from capability scores, and give the pilot a clear escalation path when inputs are ambiguous instead of rewarding an agent for guessing its way through a transaction.

Today’s HydraFusion analysis illustrates why this discipline matters even when the vendor publishes favorable cost results. Workflow selection can improve a bounded task while leaving long, changing sessions unresolved. OSWorld tests precisely those accumulating dependencies. A strong coding result is not evidence that the same system can independently reconcile an afternoon’s worth of changing business information.

The strongest counterpoint is that full automation is not every buyer’s objective. An agent that prepares a draft and hands a well-organized record to a human may be valuable without passing the benchmark’s complete-workflow test. That supports a supervised deployment, provided the handoff cost is measured. It does not justify presenting partial credit as unattended reliability.

The verdict is to upgrade the evaluation before upgrading the authority granted to the agent. Use the reported Fable progress as a reason to retest, not as a universal deployment threshold. A matched, version-pinned comparison with lower recovery costs and clean safety outcomes would support broader delegation. Failure to reproduce the gain—or gains that disappear once harmful side effects count—would leave the human at the transaction boundary. The old headline is stale; the need for that boundary is not settled by replacing it with a newer score.

Sources