Data◆Article
GPT-5.6 Sol sets a new high score of 53.6 on the Agents' Last Exam benchmark, beating Claude Fable 5 by 13.1 points, and even at medium reasoning it beats Fable 5 by 11.4 points at roughly a quarter of the cost.
OpenAI touts GPT-5.6 Sol's performance on Agents' Last Exam, a 55-field long-running professional workflow benchmark, as decisively beating Claude Fable 5, with smaller GPT-5.6 models also beating Fable 5 at much lower cost. ✦ AI generated
OpenAI · Simon Willison's Weblog · 2026-07-09 · original ↗
On Agents' Last Exam, an evaluation of long-running professional workflows across 55 fields, GPT-5.6 Sol sets a new high of 53.6, eclipsing Claude Fable 5 (adaptive reasoning) by 13.1 points. Even at medium reasoning, it beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost.
Read full article ↗excerpt · fair-use quotation
- ·New high score: 53.6 on Agents' Last Exam
- ·Benchmark spans 55 professional-workflow fields
- ·Beats Claude Fable 5 by 13.1 points
- ·Medium reasoning still beats Fable 5 by 11.4 points
- ·Medium-reasoning GPT-5.6 scores 51.9 vs Fable 5's 40.5
- ·Still an 11.4-point lead over Fable 5
- ·Costs roughly one-quarter of Fable 5's estimate
Around this claim