Contrary to claims of emergent super-intelligence, these hacking behaviors were taught: labs trained models on cybersecurity challenges with a well-defined reward function, and the novel piece is that they now also reward the path of least tokens.
Dylan asserts that labs claiming these are emergent behaviors are lying; the models were trained via reinforcement learning on well-defined cyber challenges, and the genuinely new layer is rewarding minimal-token paths, which now quantifiably reveals the path of least resistance in real security. ✦ AI generated
Dylan · a16z Podcast · 2026-08-07 · original ↗
starts at this moment · 9:27
“These are very specialized skills... What's your understanding of how they're figuring this stuff out because it seems like they've been taught to do this?”
If a lab tells you that this is an emergent super intelligence behavior, they're just lying to you. And you can read their own safety reports to see exactly how the models are trained... The interesting thing about cyber security in particular is the reward function is incredibly well defined. Get access to the data. Did it get access to the data? Reward the thing... The other piece that they've layered on top and this is where it starts to get really interesting is they've started to reward the path of least tokens. And so the reason that's interesting is because for the first time it's actually able to quantifiably show us the path of least resistance for just general cyber security to get from A to B.
verbatim transcript · starts at 9:27
9:19learned behaviors, right? And and and I think what I think we're seeing is we're seeing a a a a process that looks like it's been kind of maybe trained um or or there's a reward structure that's been built on a bunch of these things. Like what's your understanding of how they're figuring this stuff out because it it seems like they know what they're doing like they've been taught to do this.
9:40>> Yeah. I mean, if a lab tells you that this is an emergent super intelligence behavior, they're just lying to you. And you can read their own safety reports to see exactly how the models are trained and exactly how they're testing these behaviors. I mean the interesting thing about cyber security in particular is the reward function is incredibly well defined. Get access to the data. Did it
9:59get access to the data? Reward the thing. And so when they realize that like the number of problems that have that well-defined reward structure basically defines how we do reinforcement learning and they want to find as many problem spaces that they could do reinforcement learning on. And so it was a prime candidate for them to come in and give it uh CTFs and give it like cyber security challenges where
10:18they say, "Okay, get access to this thing and do whatever hacking you need to do to accomplish the goal." And then >> because they've essentially been buying pen testing data for the last four years, right? >> That's that's a piece of it. The other piece of it >> and then the capture the flag contests and all those sorts of things. >> The thing is it's just not difficult to
10:33construct a challenge. just you know even if there is no known exploit if we're talking about zero days you put a piece of software between the model and and and some data and you say get access to the data and then you know if it get access if it gets access to the data you reward it and it's that simple. But the other piece that they've layered on top
10:51and this is where it starts to get really interesting is they've started to reward the path of least tokens. And so the reason that's interesting is because for the first time it's actually able to quantifiably show us the path of least resistance for just general cyber security to get from A to B. And we've talked about like our opinions of what that is in the past of like of course uh
11:12truffle security's biased view is like a password laying around is a shorter path than going through a fancy zero day. But actually watching the model physically get from A to B and watching it follow the password and quantifying how many tokens it took to go this route versus that route. I mean, it's just incredible to watch that lay out. And it's all in their safety reports like as they um
11:33test the models out and and show, okay, well, it got access to the data and it broke out of its harness. Um it's it's not like this is emergent behavior specifically. >> That's perfectly logical, right? Like the fastest way to get a gallon of milk is to steal it. >> That's exactly right. [laughter] So, so, um, I mean, what was interesting is we were in the middle of partnering with
- ·Labs claiming emergent super-intelligence are lying
- ·Models trained via RL on well-defined cyber challenges
- ·Reward function precisely defined: 'get access to data'
- ·Claims contradicted by labs' own safety reports
- ·Novel piece rewards 'the path of least tokens'
- ·First time quantifiably showing path of least resistance
- ·Reveals shortest route from A to B in real security
- ·Interesting because resistance becomes measurable