ATRIUMsearch → argument graph
MechanismVideo · 17:45 — 19:25

AI capabilities emerge discontinuously — models jump from being unable to do something to reliably doing it at unpredictable thresholds — which makes adaptability essential and safety testing inherently harder because you cannot predict exactly when a capability will appear.

Diane Penn references the original scaling law papers, noting that while loss decreases smoothly, capabilities emerge as discontinuous jumps — a model may go from unable to calculate 1+1 to reliably doing so — and this unpredictability is why evals and safety testing are critical yet difficult. ✦ AI generated

Diane Penn · Lenny's Podcast · 2026-07-26 · original ↗

starts at this moment · 17:45

Elicited by

So what I'm hearing here is you almost don't know what will be possible with every model release. And so the important things to focus on is being adaptable as things emerge.

I think um there's some really interesting graphs in the original scaling law papers and I think folks are very familiar with the scaling loss in in the lens of um as you add in more compute and data what's called loss aka the loss from next token prediction uh goes down. And so it's a very smooth linear curve of like the models get more intelligent as you scale them up. What's actually also interesting uh in that paper is there are these like very uh different emerging capability graphs. And so for example uh as you add in more data and you train the models with more compute you essentially see these actually discontinuous emerging capabilities jump. So the models go from 1 + one being a thing that it can't calculate to a thing that it can reliably calculate. And so these emerging capabilities, this like some nature of like predictability is is is not necessarily everyone knows the exact moment like you need the evals to be able to assess that has actually always been a part of uh how this technology works and also what makes like things like safety harder because unless you have the eval unless you have the systems to test um these jumps might actually happen and you don't know

verbatim transcript · starts at 17:45

Transcript · around this moment

17:45it. So the product making it easy and even just like telling you here's something you could do feels like an important part. Is that roughly what you're describing? >> I I think so. I think um there's some really interesting graphs in the original scaling law papers and I think folks are very familiar with the scaling loss in in the lens of um as you add in more compute and data what's called loss

18:08aka the loss from next token prediction uh goes down. And so it's a very smooth linear curve of like the models get more intelligent as you scale them up. What's actually also interesting uh in that paper is there are these like very uh different emerging capability graphs. And so for example uh as you add in more data and you train the models with more compute you essentially see these

18:34actually discontinuous emerging capabilities jump. So the models go from 1 + one being a thing that it can't calculate to a thing that it can reliably calculate. And so these emerging capabilities, this like some nature of like predictability is is is not necessarily everyone knows the exact moment like you need the ebells to be able to assess that has actually always been a part of uh how this technology

19:03works and also what makes like things like safety harder because unless you have the eval unless you have the systems to test um these jumps might actually happen and you don't know M that's so interesting that you may have developed this like AI brain that uh can do something you're not even aware of and so part of the job is just uncovering wow we just got really good

19:25at this thing what can we do with that >> I think there's like product overhang and user overhang like to to maybe put it in our um PM language even on today's models and I think there's like a lot that uh we could be exploring on like our current opuses and definitely with like Fable for example temple and that that discovery is actually another part of what's been in the early days of

19:52anthropics DNA and I think is also continuing to be a big part of how we operate in product in labs and and across research. This makes me think about something Gary Tan's been talking about uh president of YC. I don't know what his title is. uh he's he had this interesting point that if you're willing to spend $100,000 a year right now in tokens, you are living the

Around this claim