Computer-use agents, the ones that see a screenshot and click like a person, have improved fast. The honest picture is messier than the demos suggest.
Our look at the OSWorld 2.0 benchmark found leading agents dropping from roughly 85% success on an easier, saturated test to around 20% on a harder one testing longer task chains. Any “best computer-use agent” claim citing a number is only meaningful once you know which version of the benchmark it’s from.
They’re genuinely good at short, well-defined tasks with a clear visual target, and at software with no usable API at all. They’re still bad at long workflows spanning multiple applications, unfamiliar layouts, and recovering from an unexpected pop-up.
Test on your actual longest, messiest task, not a vendor demo. And always prefer a direct API integration when one exists, screen automation should be the fallback, not the first choice.
Review the benchmark directly at the OSWorld project page.




