DataArticle · 7:24 · 2m
Kimi K3 produces significantly simpler, more readable code than Codex, but misses edge cases that Codex catches, making it roughly equivalent to a GPT-54 or 55 level for coding tasks.
Florian Brand · Interconnects
Despite being difficult, games in DiG-bench are beatable by humans, but today's best AI models struggle significantly, indicating a gap in discovery capabilities. ✦ AI generated
author · Import AI · 2026-08-17 · original ↗
Overall, this seems really hard! ... some frontier models are already capable of some fairly impressive feats of discovery, but still struggle compared to humans (for instance, a 20% success rate on Tier 7 is pretty poor compared to the fact individual humans were able to get 100% on the tests).
Read full article ↗excerpt · fair-use quotation