Data◆Article · 47 words
GPT-5.6 Sol sets a new high score of 53.6 on the Agents' Last Exam benchmark, beating Claude Fable 5 by 13.1 points, and even at medium reasoning it beats Fable 5 by 11.4 points at roughly a quarter of the cost.
OpenAI · Simon Willison's Weblog