Claim◆Article · 1:17 — 2:47
The number of months open models are behind closed frontier models is impossible to pin down because different benchmarks give different answers.
Florian argues that measuring the open-closed gap in months is meaningless because benchmark providers cherry-pick metrics to support their preferred narrative. ✦ AI generated
Florian Brand · Interconnects · 2026-07-22 · original ↗
The biggest thing at every model release, at least at every open model release, is how much or how many months it is behind the closed frontier. People love to put a definite number onto this, which is really muddying because we have so many different benchmark providers and such so many different benchmarks that every site pulls up their favorite benchmarks to show that the current model is at the frontier, which is then countered by the other side pulling up another benchmark and showing it's actually a year behind.
Read full article ↗excerpt · fair-use quotation
- ·Every open model release starts a months-behind debate
- ·Benchmark providers cherry-pick metrics to fit their narrative
- ·No single number can measure the gap objectively
- ·One side shows a model at the frontier using its favorite benchmark
- ·The other side cites a different benchmark to claim it's a year behind
- ·The result is muddying, not clarifying
Around this claim
This moment responds to
rebuts → The performance gap—both open-to-closed and American-to-Chinese—has shrunk from a debated 6-9 months down to something like 3-5 months.Nathan Lambert · Interconnectsgives example → The US cybersecurity community was forced to use a weaker Chinese model (GLM) to analyze a hacking attempt because frontier models had guardrails blocking defensive analysis.Florian Brand · Interconnectsprovides context → When reviewing a project, benchmarks are unsatisfying because napkin math shows you what the theoretical lower bound should be, and a large gap means either the benchmark is wrong or your understanding is flawed.Simon Eskildsen · The Pragmatic Engineer