AI agents just beat humans (72%) at using a computer. 85%.

But a harder version of the exact same benchmark only got 20.6%.

The numbers didn't lie. We just never asked what the conditions were.

3-second summary
85% benchmark beats humans (72%) but OSWorld 2.0 says 20.6% why: different measurement 5-point vendor checklist

Here's what everyone believes

a16z's report last week didn't hedge. The title was "Can Agents Use a Computer Yet? We've Got the Data." The conclusion: yes, it's here.

85%
OSWorld-Verified, Claude Fable 5
72%
Human tester baseline
12%→85%
April 2024 → June 2026

Two years ago, the top model on OSWorld-Verified scored 12%. Now Claude Fable 5 hits 85%, clearing the human baseline of 72%. a16z's Serafini, Amble, and Zhou backed it up with real math: a computer-use agent runs $6-8/hour, while an Indian BPO worker costs ~$10/hour, and a US back-office worker runs $30-45/hour once you add benefits on top of the $20.59/hour median wage.

And there are live deployments to point to. One CPG data platform runs 15-21 million automated portal interactions a month. A global systems integrator handles 1,500-2,100 IT tickets a day across 27 live workflows. At that point, "should we adopt this?" starts to feel like the obvious question.

But the number was only half the story

Around the same time, a completely different number came out of a benchmark with the exact same name. XLang Lab, the team behind OSWorld, released OSWorld 2.0 in mid-2026 — and the best-performing agent on it only completed 20.6% of tasks.

Why the gap? OSWorld 2.0 is built from 108 real tasks that take a skilled human an average of 1.6 hours to finish, spanning research, content creation, software development, business/finance and more. Each task carries an average of 27.25 checkpoints for partial credit. The original OSWorld (1.0) was mostly short, self-contained tasks by comparison.

"The gap between those two numbers is the most important fact in the field."

— Adnan Masood, AI researcher

Dev blog youngju.dev nailed the specifics: the same model can swing 63 percentage points (83.5% vs. 20.6%) on what's called the same benchmark, and the model itself didn't get worse — the measurement did. Short, self-contained tasks score in the 80s. Real, long workflows drop to the 20s. About 45% of tasks are actually solved through terminal/scripting instead of GUI clicks, which means "computer-use capability" is really a blend of GUI manipulation, scripting, and reasoning. For reference, OpenAI's Operator scored 38% on the original OSWorld — described as "close to random on many task categories".

How benchmark numbers get made

The name "OSWorld" covers both a short-task version (1.0) and a 1.6-hour-real-work version (2.0). A number without a disclosed step budget, retry count, and benchmark version isn't comparable to anything.

So what should you actually believe?

Reread a16z's report and the authors clearly knew about this trap. Their real claim wasn't "AI can now do anything on a computer" — it was the narrower one: "computer-use agents are strongest on standardized, repeatable tasks." Every deployment they cited — CRM updates, IT ticket triage, government/insurance portal logins — had a clear completion condition. A 15% failure rate is fine when there's an escalation path for a human, because 24/7 throughput makes up for it.

85% and 20.6% are answers to two different questions. What matters is which bucket your own workflow falls into.

Open-ended workStructured, repeatable work
ExamplesOpen-ended research, judgment calls, exception-heavy tasksCRM data entry, IT ticket triage, portal logins/forms
Completion conditionJudgment call every timeClear and repeatable
Benchmark evidence~20.6% on OSWorld 2.0~85% on OSWorld-Verified
Current stateHumans still need to leadProduction deployments already exist

What to ask when a vendor pitches you

  1. Ask which benchmark version first
    "We ran OSWorld" isn't enough. Ask if it was 1.0 or 2.0, and how long the average task actually took.
  2. Check the step budget and retry count
    How many retries were allowed, and how many times a human stepped in on failure — this alone can flip a score.
  3. Classify your own workflow
    Does it have a clear, repeatable completion condition, or does it require a judgment call every time?
  4. Pilot the structured work first
    Start narrow, like a16z's examples — ticket triage, data entry, portal logins.
  5. Plan for the failure rate
    If there's no escalation path for the ~15% that fails, you're not ready to deploy yet.

Want to go deeper?

Can Agents Use a Computer Yet? a16z's original report, with the full cost math and deployment cases. a16z.com

OSWorld's official benchmark site See the task composition and live leaderboard yourself. osworld-v1.xlang.ai

The Hardest Easy Problem in AI A clear breakdown of the gap between benchmark scores and real-world completion. medium.com

What Computer-Use Benchmarks Actually Measure A Korean dev blog's deep dive into what benchmark numbers really capture. youngju.dev

Computer-use agents hit 85% on OSWorld A two-year timeline of the benchmark's rapid climb. cryptobriefing.com

Korean coverage of OSWorld 2.0 A news report on the 108-task structure and scoring method. aitimes.com