ATRIUMsearch → argument graph
DataArticle

GPT-5.6 Terra performs just above Claude Fable 5 and Luna outperforms Opus 4.8, each in roughly one-third the time, half the output tokens, and about one-quarter the cost, with new state-of-the-art results on Terminal-Bench 2.1 and DeepSWE.

OpenAI claims its new Terra and Luna models beat Claude's Fable 5 and Opus 4.8 respectively, hitting new state-of-the-art scores on Terminal-Bench 2.1 and DeepSWE at a fraction of the time, tokens, and cost. ✦ AI generated

OpenAI · Latent Space · 2026-07-10 · original ↗

Terra performs just above Fable 5, while Luna outperforms Opus 4.8; each does so in roughly one-third of the time, with about half as many output tokens, and at approximately one-quarter the estimated cost. It also sets new state-of-the-art results on Terminal‑Bench 2.1 and DeepSWE, which test complex command-line workflows and long-horizon engineering in real codebases.

Read full article ↗excerpt · fair-use quotation

Around this claim