DataArticle
Current frontier AI agents remain far from reliable at long-horizon, multi-step computer-use tasks, with even the strongest configuration (Claude Opus 4.8) scoring only 20.6% binary accuracy on OSWorld 2.0, struggling especially with hidden-state recovery, tracking many items, and conflicting information.
OSWorld 2.0, a benchmark of 108 long-horizon computer-use tasks averaging 1.6 hours for a human to complete, shows even top models like Claude Opus 4.8 achieving only 20.6% binary accuracy, though rapid gains are expected as happened with OSWorld 1.0. ✦ AI generated
OSWorld 2.0 paper authors · Import AI · 2026-07-06 · original ↗
Our experiments show that current agents remain far from reliable computer use: the strongest setting, Claude Opus 4.8 with maximum thinking and batched tool calls, reaches only 20.6% binary accuracy and 54.8% partial-score accuracy, they write. Performance drops sharply as tasks grow longer, and agents struggle most when they must recover hidden state, track many items, resolve conflicting information, or adapt to changing requirements.
Read full article ↗excerpt · fair-use quotation
Around this claim
Evidence · 2
Compounding error multiplies across chained steps: if each model call is correct 95% of the time, running twenty steps in sequence succeeds only about one in three times.Best Practices for Building AI Agents That Work in Production (author) · ByteByteGo Newsletter · conf 85%Achieving human-level power efficiency in robotics for menial tasks will take a very long time; the system complexity—hardware, software, brain, even finger friction coefficients—demands extensive iteration, unlike LLMs where performance-to-power is already close in narrow tasks.Yunu Guo & Fei-Fei Li · a16z Podcast · conf 75%
This moment responds to
supports → Current frontier AI agents remain far from reliable at long-horizon, multi-program computer-use tasks, with the best-performing setup reaching only 20.6% binary accuracy on OSWORLD 2.0, though rapid improvement is likely given how quickly scores rose on OSWORLD 1.0.OSWorld 2.0 researchers · Import AIextends → Current frontier AI agents remain far from reliable at long-horizon, multi-program computer-use tasks, with the best-performing setup reaching only 20.6% binary accuracy on OSWORLD 2.0, though rapid improvement is likely given how quickly scores rose on OSWORLD 1.0.OSWorld 2.0 researchers · Import AIrebuts → Using Priceline's AI agent 'Penny,' Fogel personally planned a complex multi-city family trip — splitting cabins, mixing miles and cash, and sequencing hotel and onward travel — with the AI asking clarifying questions and working through tradeoffs interactively.Glenn Fogel · No Priors