Computer-Use Agents Just Went From 85% to 20% on a Harder Benchmark

Computer-use agents just went from bragging rights to a reality check in a single release cycle. The field’s flagship benchmark, OSWorld, had become so thoroughly beaten by frontier models that scores were pushing past 85% — comfortably ahead of the roughly 72% human baseline. Then on June 26, 2026, the same research group released a harder version, OSWorld 2.0, and the best model in the world dropped straight back down to around 20.6%.

Quick facts

  • OSWorld (the original benchmark) measures how well AI agents complete real desktop tasks — 369 tasks spanning web and desktop apps, file operations, and multi-app workflows — using only screenshots, clicks, and keystrokes, the same interface a human uses.
  • Top computer-use agents reached roughly 85% on OSWorld by mid-2026, up from around 12% in April 2024 and ahead of the ~72% human baseline.
  • OSWorld 2.0, released June 26, 2026, is designed to capture the realism, complexity, and long-horizon demands the original benchmark missed.
  • On the new benchmark, the best-performing model’s score fell to roughly 20.6% — a genuine capability gap, not a rounding difference.

Why the old benchmark stopped meaning much

Benchmark saturation is a familiar pattern in AI: a test is hard, models improve until they beat it, and eventually the test stops distinguishing genuinely capable systems from ones that have simply learned the test’s specific patterns. OSWorld followed that arc unusually fast — from roughly 12% success in April 2024 to around 85% little more than two years later, according to independent analysis published on Medium. That’s an unusually steep improvement curve for a benchmark meant to represent real-world computer use, and it’s exactly the kind of trajectory that should make anyone skeptical that the underlying capability actually improved as fast as the score did.

What OSWorld 2.0 actually tests differently

The research team behind the original benchmark built OSWorld 2.0 specifically because, in their own framing, existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of how people actually use computers — limiting the benchmark’s ability to reveal where frontier agents genuinely fall short. Rather than adjusting scoring on the same task set, the new benchmark changes what’s being asked: longer task chains, more realistic multi-step workflows, and less tolerance for the kind of narrow pattern-matching that let agents rack up points on the original test without a durable underlying capability.

The result, per an independent field guide covering the benchmark shift, was a genuine inversion rather than a gradual correction: the best model available dropped from beating the human baseline to completing roughly one in five tasks, in a single release. The guide’s framing is blunt about what that means for anything written about computer-use agents before this benchmark existed — most of the 2025-era leaderboard claims are now effectively obsolete.

What this means if you’re evaluating a computer-use agent

The practical lesson isn’t that computer-use agents are useless — it’s that a headline benchmark score tells you almost nothing without knowing which benchmark, and how recently it was set. An 85% OSWorld score earned before June 2026 and a 20% OSWorld 2.0 score earned after it can both be accurate descriptions of the same underlying model on different tests. Treat any vendor’s computer-use benchmark claim as incomplete until you know which version of which benchmark it’s citing, and be specifically wary of vendors quoting only their best number without naming the test.

For real deployments, the gap between short, narrow tasks and long, multi-step workflows is exactly where computer-use agents still struggle most — consistent with what OSWorld 2.0 is specifically designed to expose. If you’re evaluating a computer-use agent for a real workflow, test it on your actual longest, messiest task, not a demo-friendly short one.

Common questions

Is OSWorld 2.0 harder for every model equally? The reporting reviewed here doesn’t provide a full model-by-model breakdown; the headline figure is the best available model’s score falling to roughly 20.6%, which implies every model dropped substantially, but relative rankings between specific models on the new benchmark weren’t detailed in the sources available.

Does this mean computer-use agents got worse? No — the models didn’t change. What changed is the test. The same agents that scored 85% on the original OSWorld are the ones scoring 20.6% on OSWorld 2.0; the underlying capability is the same, but the harder benchmark reveals limitations the easier one didn’t test for.

Should I stop trusting OSWorld scores? Not entirely, but always check which version. “OSWorld” without a version number, especially in vendor marketing, should be treated as a yellow flag until you confirm whether it’s the original benchmark or OSWorld 2.0 — the two numbers aren’t comparable.

Key takeaway

Computer-use agents genuinely improved enormously between 2024 and mid-2026 — the 12%-to-85% climb on the original OSWorld is real. But OSWorld 2.0’s results are a useful corrective against treating that climb as evidence the problem is solved: on harder, more realistic tasks, the field is closer to where it started than the old leaderboard suggested.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *