Anthropic's internal Model 2 scores significantly higher on internal benchmarks than any publicly released model, creating a dangerous and widening gap between what companies have internally and what the public can access.
Dylan Patel of SemiAnalysis leaked that Anthropic has finished training Mythos 2 (called 'Model 2' internally) and is not releasing it publicly. Corroborated by Anthropic's own redacted risk report, Model 2 scores 1.5 points higher on the Epoch Capabilities Index and 8.8 percentage points higher on Co-Bench than Mythos Preview. The gap between internal and public capabilities is growing rapidly. ✦ AI generated
Pash · The Cognitive Revolution · 2026-08-18 · original ↗
starts at this moment · 1:42
Anthropic has Mythos 2 which has finished training and that they're not releasing it to the public essentially... Model 2 is an internal model that they do not plan to release to the public. It scores 1.5 what they call ECI, AECIS which is the epoch capabilities index. So it scores 1.5 points higher than their previous high model... On Co-Bench, Anthropic's Model 2 is about 8.8 percentage points higher than Mythos Preview... Mythos Preview itself is about 4 percentage points higher than Mythos 5. And Mythos 5 was actually almost double of Claude Opus 4.7.
verbatim transcript · starts at 1:42
1:42know how um how accurate or how relevant it is but it comes from our friend uh Dylan Patel who is the uh editor of semi analysis. I'm going to play this clip. So this is on the AI spend and this is on ethos. >> I don't hear it. It's on It's on the stream. So for those of you who are kind of looking at it, so Dylan comes out and says um
3:06Anthropic has mythos 2 which has finished training and um that they're not releasing it to the public essentially. And that's where um I think he stands there. And the question is of course how reliable a source Dylan is because he is you know an editor editor at large. Um and so we we we do have some questions on how reliable as a source he is. But we do
3:36also have to note that number one Dylan is u flatmates with Shelter Douglas at Entropic. He is best friends with um Leopold Ashen Brener who is married to Avital Bowit who is the chief of staff at Entropic. So we have these um many of these links in there. Dylan is deeply embedded in the community. I don't think you know anyone in the community is leaking to him but engineers
4:07>> never never >> never could possibly >> never never there should there'll never be any of these leaks but engineers especially research engineers um often regard their work as um not particularly um you know unless you win a fields medal for it it's not exactly like the most technical stuff that you're doing you I I know externally a lot of people think it's like super technical and
4:34super difficult, but for the guy who's been doing like you know GPU kernels for like 10 years, another GPU kernel is not exactly like the most um you know innovative or intellectual thing. And so they often are uh you know quite free uh you know when you when you meet with them after hours. Uh and so I think it's fairly credible that you know they he he
4:55said that uh I will say um and I'm going to I'm going to share another um so uh anthropic put out so they didn't they put out this report. It's a redacted risk report. Um let me just pull it up on screen here. Um if we can get that. Yes, there we go. Um, and so this is their redacted risk report. They put this out uh a couple of
5:24weeks ago. Um, and they do talk about model 2. It's not mythos 2, they call it model 2. Uh, model 2 is an internal model that they do not plan to release to the public. Um it scores 1.5 what they call ECI AECIS which is the epoch capabilities index. Uh so it scores 1.5 points higher uh than their previous high model. That is roughly a six- week kind of gap
5:581.5 points uh if you if you look at their at the at the metrics that they they've given. But they hide elsewhere in the report this co-bench score. And co-bench is a uh metric of internal anthropic uh research problems and how good the how how well the models actually uh accelerate or help them uh on these uh on these metrics. And on cobbench uh enthropic uh anthropics
- ·Mythos 2 (Model 2) finished training — not releasing publicly
- ·Scores higher than any released model on internal benchmarks
- ·Internal-public capability gap is widening rapidly
- ·Model 2 beats Mythos Preview by 8.8 percentage points
- ·Mythos Preview itself exceeds Mythos 5 by ~4pp