ATRIUMsearch → argument graph
DataArticle

GPT-5.6 Sol sets a new high score of 53.6 on the Agents' Last Exam benchmark, beating Claude Fable 5 by 13.1 points, and even at medium reasoning it beats Fable 5 by 11.4 points at roughly a quarter of the cost.

OpenAI touts GPT-5.6 Sol's performance on Agents' Last Exam, a 55-field long-running professional workflow benchmark, as decisively beating Claude Fable 5, with smaller GPT-5.6 models also beating Fable 5 at much lower cost. ✦ AI generated

OpenAI · Simon Willison's Weblog · 2026-07-09 · original ↗

On Agents' Last Exam, an evaluation of long-running professional workflows across 55 fields, GPT-5.6 Sol sets a new high of 53.6, eclipsing Claude Fable 5 (adaptive reasoning) by 13.1 points. Even at medium reasoning, it beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost.

Read full article ↗excerpt · fair-use quotation

Around this claim