ATRIUMsearch → argument graph
ClaimArticle

Claude Fable 5 beat GPT-5.6 Sol badly on SWE-Bench Pro (80% vs 64.6%), a result OpenAI downplays by claiming roughly 30% of SWE-bench Pro tasks are broken and urging developers to scrutinize results carefully.

Fable 5 crushed the GPT-5.6 family on SWE-Bench Pro, and OpenAI published a separate audit the day before claiming ~30% of that benchmark's tasks are broken, casting doubt on the result. ✦ AI generated

OpenAI · Simon Willison's Weblog · 2026-07-09 · original ↗

In light of these results, we estimate that ~30% of SWE-bench Pro tasks are broken, and advise that model developers carefully examine results

Read full article ↗excerpt · fair-use quotation

Related moments