ATRIUMsearch → argument graph
Article · 2026-07-10 · 6 moments

[AINews] OpenAI launches GPT 5.6 Sol/Terra/Luna, Codex becomes ChatGPT superapp

A big day for OpenAI. ✦ AI generated

01
Mechanism

GPT-5.6's new ultra effort level coordinates four agents in parallel by default, trading higher token use for stronger results and faster time-to-result on demanding tasks.

OpenAI describes a new 'ultra' reasoning setting beyond 'max' that runs four agents in parallel by default, using more tokens to get stronger, faster results on hard tasks.

transcript

OpenAI: max gives GPT‑5.6 even more time than xhigh to reason and explore alternatives, run checks, and revise its approach. ultra goes further by coordinating four agents in parallel by default, trading higher token use for stronger results and faster time-to-result on demanding tasks.

02
Claim

GPT-5.6 Sol autonomously post-trained GPT-5.6 Luna.

OpenAI's launch materials stated that its flagship Sol model autonomously post-trained the smaller Luna model, a claim that spread quickly and fueled speculation about automated AI research.

transcript

OpenAI: Multiple accounts amplified the statement that OpenAI says GPT‑5.6 Sol autonomously post-trained GPT‑5.6 Luna, via @scaling01, @tejalpatwardhan, and @dejavucoder. The claim fueled RSI/autoresearch speculation; @tenobrus said if true as stated, it would be a "pretty large update" for automated researcher timelines.

rebuts · 1

03
Data

GPT-5.6 Terra performs just above Claude Fable 5 and Luna outperforms Opus 4.8, each in roughly one-third the time, half the output tokens, and about one-quarter the cost, with new state-of-the-art results on Terminal-Bench 2.1 and DeepSWE.

OpenAI claims its new Terra and Luna models beat Claude's Fable 5 and Opus 4.8 respectively, hitting new state-of-the-art scores on Terminal-Bench 2.1 and DeepSWE at a fraction of the time, tokens, and cost.

transcript

OpenAI: Terra performs just above Fable 5, while Luna outperforms Opus 4.8; each does so in roughly one-third of the time, with about half as many output tokens, and at approximately one-quarter the estimated cost. It also sets new state-of-the-art results on Terminal‑Bench 2.1 and DeepSWE, which test complex command-line workflows and long-horizon engineering in real codebases.

provides context · 1

04
Claim

Sol did not autonomously conduct end-to-end post-training or research on Luna; what likely happened is a model implementing small graders, reward-shaping logic, or training configs on top of OpenAI's existing mature RL infrastructure.

Skeptic scaling01 pushed back on the viral 'Sol post-trained Luna' claim, arguing it reflects the model editing small configs and reward-shaping logic within mature infrastructure, not genuine autonomous end-to-end research.

transcript

scaling01 (X/Twitter): @scaling01 argued that what's probably happening is a model implementing LLM-as-a-judge graders, reward-shaping logic, or small training configs on top of existing OpenAI RL infrastructure—not autonomous end-to-end research or training systems. @scaling01 explicitly said we should distance these statements from literal autonomous end-to-end post-training or research, which models still cannot do.

provides context · 1

05
Fact

Testing found universal jailbreaks in GPT-5.6 Sol across every round that enabled long-form agentic task completion in vulnerability discovery and exploit development, making it the highest-stakes safety issue of any model release yet.

An AI Safety Institute researcher reported finding universal jailbreaks in every round of testing GPT-5.6 Sol, unlocking long-form agentic vulnerability discovery and exploit development, which a colleague called the highest-stakes safety issue of any model release yet.

transcript

alxndrdavies (AI Safety Institute): @alxndrdavies from the AI Safety Institute said they found universal jailbreaks in all rounds of testing that enabled long-form agentic task completion in vulnerability discovery and exploit development. @EthanJPerez called it "the highest stakes safety issue of any model release yet"

rebuts · 2

06
Context

The bundling of GPT-5.6 with ChatGPT Work, Sites, and the Codex desktop merge shows OpenAI moving from a model vendor into a full-stack work platform, with researchers already using these systems to automate chunks of RL/post-training workflows.

The newsletter's analysis argues that bundling GPT-5.6 with ChatGPT Work, a merged Codex/ChatGPT desktop app, and Sites hosting signals OpenAI's strategic shift from selling models to owning the whole agentic work stack, with internal researchers already automating parts of model-improvement work.

transcript

AI News (Latent Space): The product bundling suggests OpenAI is moving from a model vendor to a full-stack work platform, with its own browser, connectors, orchestration primitives, hosted app deployment, and desktop runtime. The strongest forward-looking signal may be the internal claim that researchers already use these systems to materially increase output and automate chunks of RL/post-training workflows, even if public discussion often overstates that as "the model trained itself"

extends · 3rebuts · 1

Highlight slides
OpenAI: Sol Autonomously Post-Trained Luna✦ from: GPT-5.6 Sol autonomously post-trained GPT-5.6 Luna.OpenAI Claims Wins Over Claude Models✦ from: GPT-5.6 Terra performs just above Claude Fable 5 and Luna outperforms Opus 4.8, each in roughly one-third the time, half the output tokens, and about one-quarter the cost, with new state-of-the-art results on Terminal-Bench 2.1 and DeepSWE.Universal Jailbreaks Found in GPT-5.6 Sol✦ from: Testing found universal jailbreaks in GPT-5.6 Sol across every round that enabled long-form agentic task completion in vulnerability discovery and exploit development, making it the highest-stakes safety issue of any model release yet.Highest-Stakes Safety Issue Yet✦ from: Testing found universal jailbreaks in GPT-5.6 Sol across every round that enabled long-form agentic task completion in vulnerability discovery and exploit development, making it the highest-stakes safety issue of any model release yet.Efficiency vs. Claude Models✦ from: GPT-5.6 Terra performs just above Claude Fable 5 and Luna outperforms Opus 4.8, each in roughly one-third the time, half the output tokens, and about one-quarter the cost, with new state-of-the-art results on Terminal-Bench 2.1 and DeepSWE.Claim Spreads Fast, Fuels RSI Speculation✦ from: GPT-5.6 Sol autonomously post-trained GPT-5.6 Luna.Skepticism on What It Would Mean✦ from: GPT-5.6 Sol autonomously post-trained GPT-5.6 Luna.New State-of-the-Art Results✦ from: GPT-5.6 Terra performs just above Claude Fable 5 and Luna outperforms Opus 4.8, each in roughly one-third the time, half the output tokens, and about one-quarter the cost, with new state-of-the-art results on Terminal-Bench 2.1 and DeepSWE.
Related episodes