Frontier Code is designed to judge whether AI-written code would actually be merged by a human reviewer, not just whether it passes tests, because saturated benchmarks like SWE-bench allow heavy reward-hacking and cheating.
swyx explains that Cognition built Frontier Code as an out-of-sample, heavily rubricked benchmark because saturated benchmarks like SWE-bench let models cheat their way to passing scores while producing code that's often unmergeable. ✦ AI generated
swyx · The Cognitive Revolution · 2026-06-27 · original ↗
starts at this moment · 53:05
“We asked what makes it different from the benchmarks that came before.”
The reason that we were so excited about Frontier Code is it, you know, you stop being able to articulate the differences in model quality with more saturated benchmarks like Swebench because they're all like at most you'll get like a 1 to 2% bump and they're like cool, but like how much of that is memorization or what have you. Frontier Code is all out of sample.
verbatim transcript · starts at 53:05
53:05different from the benchmarks that came before. [music] The reason that we were so excited about Frontier Code is is it, you know, you stop being able to articulate the differences in model quality with more saturated benchmarks like Swebench because they're all like at most you'll get like a 1 to 2% bump and they're like cool, but like how much of that is memorization or what have you. Um,
53:32Frontier Code is all out of sample. They're not in the training set. They're all graded and heavily rubricked like we basically found like deep sweet frontier code sorry deep sweet bbench all these things they actually allow a lot of false positives in the way that models can cheat in the same way that uh you know during training they they basically have reward hacks same thing and we have
53:56an internal catalog of like 20 different ways that models cheat and so basically we just translated that to your rubrics and you know ship ship that as one to your code and I think like that that is like how we want to judge models going forward. Um the not just that whether they can pass the test but can they write code that we would merge right? Meter had this very very interesting
54:15blog post where they were like about 50% of Sweepbench code that passes the Sweetbench test is completely unmergable. Like it's it's like so low quality like yeah like technically you'll pass but like you know like just on really stupid benchmarks like did you like modify a whole bunch of files you weren't supposed to touch or like uh did you cheat on the test or did you adhere
54:37to like code style what have you? uh just completely unverirtual and so like yeah we we want to guide the evolution of models towards maintainable code and against slob [music] >> then pash asked when frontier code itself gets saturated. So, so there's two parts of the strategy. Front, uh, Frontier Code 2026 will be saturated by the end of this year. We, you know, my estimate that you, we'll probably hit like 80% by the
55:05end of this year. That's as designed, that's expected. Um, it is based on open- source repos, uh, which will eventually get trained on. So, they they just leak. So like you're screwed if you think if you want one uh benchmark that will never get saturated if you're especially if you're based in open source. So the answer is very simple just do annual cadences. So then we'll have frontier code 2027 2028 all these