Anthropic's internal Model 2 scores significantly higher on internal benchmarks than any publicly released model, creating a dangerous and widening gap between what companies have internally and what the public can access.
Dylan Patel of SemiAnalysis leaked that Anthropic has finished training Mythos 2 (called 'Model 2' internally) and is not releasing it publicly. Corroborated by Anthropic's own redacted risk report, Model 2 scores 1.5 points higher on the Epoch Capabilities Index and 8.8 percentage points higher on Co-Bench than Mythos Preview. The gap between internal and public capabilities is growing rapidly.
transcript
Pash: Anthropic has Mythos 2 which has finished training and that they're not releasing it to the public essentially... Model 2 is an internal model that they do not plan to release to the public. It scores 1.5 what they call ECI, AECIS which is the epoch capabilities index. So it scores 1.5 points higher than their previous high model... On Co-Bench, Anthropic's Model 2 is about 8.8 percentage points higher than Mythos Preview... Mythos Preview itself is about 4 percentage points higher than Mythos 5. And Mythos 5 was actually almost double of Claude Opus 4.7.