DataArticle
Current frontier AI agents remain far from reliable at long-horizon, multi-program computer-use tasks, with the best-performing setup reaching only 20.6% binary accuracy on OSWORLD 2.0, though rapid improvement is likely given how quickly scores rose on OSWORLD 1.0.
OSWorld 2.0, a benchmark of multi-hour, multi-program computer-use tasks, shows even Claude Opus 4.8 with maximum thinking only reaches 20.6% binary accuracy, though the field expects a rapid ramp similar to OSWORLD 1.0's rise from ~30% to ~75%. ✦ AI generated
OSWorld 2.0 researchers · Import AI · 2026-07-06 · original ↗
Our experiments show that current agents remain far from reliable computer use: the strongest setting, Claude Opus 4.8 with maximum thinking and batched tool calls, reaches only 20.6% binary accuracy and 54.8% partial-score accuracy," they write.
Read full article ↗excerpt · fair-use quotation
Around this claim
This moment responds to
rebuts → By any definition used 20 or more years ago, we have already achieved artificial general intelligence, even though it hasn't been fully deployed yet.Jason Calacanis · All-In Podcastsupports → AI doesn't actually make work easier, and using it should force you to think just as hard or harder — so be deliberate about which skills you're willing to let atrophy.Gergely Orosz · The Pragmatic Engineerrebuts → Qwen3.8-Max demonstrates autonomous long-horizon and multimodal capabilities including 10+ days of unattended coding, a 125-hour autonomous research loop that beat an original paper's benchmark by +2.71 points, silicon design from RTL to physical layout, and product-grade outputs across professional workflows.AINews · Latent Space