Agents that operate a computer the way a person does went from research curiosity to a genuinely tracked benchmark category remarkably fast, and the pace of that progress is part of the story.
Before general-purpose computer-use agents, automating a UI meant scripted, brittle tools tied to a specific application’s exact layout, breaking the moment that layout changed.
Once models could reliably interpret a screenshot and identify interactive elements without hardcoded rules, the category became genuinely general-purpose: an agent could attempt to operate unfamiliar software the way a person encountering it for the first time would, at least in principle.
Our coverage of the OSWorld 2.0 benchmark covers the honest version of this: scores on the original test climbed from roughly 12% to 85% in about two years, a curve steep enough to reflect the benchmark becoming saturated, not the underlying problem actually being solved. A harder version brought leading agents back down to around 20%.
Short, well-defined tasks and software with no usable API are genuinely reliable use cases now. Long, multi-app workflows with unfamiliar layouts remain the frontier. Review the benchmark directly at the OSWorld project page.




