ATRIUMsearch → argument graph
DataArticle

Current frontier models cannot beat DiG-bench, with only Opus 5 and Fable 5 reaching Tier 7, suggesting AI still struggles with creative discovery compared to humans.

Despite being difficult, games in DiG-bench are beatable by humans, but today's best AI models struggle significantly, indicating a gap in discovery capabilities. ✦ AI generated

author · Import AI · 2026-08-17 · original ↗

Overall, this seems really hard! ... some frontier models are already capable of some fairly impressive feats of discovery, but still struggle compared to humans (for instance, a 20% success rate on Tier 7 is pretty poor compared to the fact individual humans were able to get 100% on the tests).

Read full article ↗excerpt · fair-use quotation

Related moments