There are discontinuous emerging capability jumps as you scale models — the models go from being unable to calculate something to reliably calculating it. These jumps are not perfectly predictable; you need the evals and systems to test for them, and that unpredictability is also what makes safety harder.
Diane explains that scaling laws produce smooth loss curves but discontinuous jumps in emerging capabilities — models unpredictably gain new abilities. This unpredictability is core to how the technology works and makes safety testing essential because you might not know a new capability exists until you test for it. ✦ AI generated
Diane Penn · Lenny's Podcast · 2026-07-26 · original ↗
starts at this moment · 17:46
“So what I'm hearing here is you almost don't know what will be possible with every model release. And so the important things to focus on is being adaptable as things emerge.”
There's some really interesting graphs in the original scaling law papers. Folks are very familiar with the scaling loss in the lens of as you add in more compute and data, the loss from next token prediction goes down. So it's a very smooth linear curve of the models get more intelligent as you scale them up. What's actually also interesting in that paper is there are these very different emerging capability graphs. For example, as you add in more data and you train the models with more compute, you see these actually discontinuous emerging capabilities jump. So the models go from 1 + 1 being a thing that it can't calculate to a thing that it can reliably calculate. And so these emerging capabilities, this nature of predictability is not necessarily everyone knows the exact moment — you need the evals to be able to assess that — has actually always been a part of how this technology works and also what makes things like safety harder because unless you have the eval, unless you have the systems to test, these jumps might actually happen and you don't know.
verbatim transcript · starts at 17:46
17:45it. So the product making it easy and even just like telling you here's something you could do feels like an important part. Is that roughly what you're describing? >> I I think so. I think um there's some really interesting graphs in the original scaling law papers and I think folks are very familiar with the scaling loss in in the lens of um as you add in more compute and data what's called loss
18:08aka the loss from next token prediction uh goes down. And so it's a very smooth linear curve of like the models get more intelligent as you scale them up. What's actually also interesting uh in that paper is there are these like very uh different emerging capability graphs. And so for example uh as you add in more data and you train the models with more compute you essentially see these
18:34actually discontinuous emerging capabilities jump. So the models go from 1 + one being a thing that it can't calculate to a thing that it can reliably calculate. And so these emerging capabilities, this like some nature of like predictability is is is not necessarily everyone knows the exact moment like you need the ebells to be able to assess that has actually always been a part of uh how this technology
19:03works and also what makes like things like safety harder because unless you have the eval unless you have the systems to test um these jumps might actually happen and you don't know M that's so interesting that you may have developed this like AI brain that uh can do something you're not even aware of and so part of the job is just uncovering wow we just got really good
19:25at this thing what can we do with that >> I think there's like product overhang and user overhang like to to maybe put it in our um PM language even on today's models and I think there's like a lot that uh we could be exploring on like our current opuses and definitely with like Fable for example temple and that that discovery is actually another part of what's been in the early days of
- ·Loss curves are smooth; capability curves are not
- ·Models can jump from inability to reliable performance
- ·Jump timing is unpredictable without evals
- ·Unpredictability is inherent to the technology, not a bug
- ·Need evals and systems to detect new capabilities
- ·Unpredictable jumps make safety harder
- ·Capabilities may exist before you test for them
- ·Scaling alone doesn't reveal what models can do