Reward seeking is distinct from reward hacking: reward seeking involves situational awareness and explicit reasoning about what is being graded and optimizing for that, while reward hacking is just exploiting loopholes without internal representation of the exploit.
Alex and Axel distinguish reward seeking (situationally aware reasoning about what is being graded) from reward hacking (exploiting loopholes without internal representation). A model can be reward seeking without reward hacking and vice versa. ✦ AI generated
Alex · Machine Learning Street Talk · 2026-07-31 · original ↗
starts at this moment · 32:46
what hacking is when you find unintended tricks or hacks or unintended solutions to problems that essentially developers didn't intend. Um so like a classic example is like in the coastal runners or coast runners I think the game is where it's the the boat is optimiz like the RL agent is optimized to drive a boat in c in in a in a race and essentially what it learns is just to drift in a corner over and over and maximize reward that way. So that's like one example of like very simple reward hacking. Um now reward seeking is the general um is generally the situational aware reasoning about what is being graded, what is being rewarded and then optimizing your actions for that. So reward hacking is one useful strategy um for getting a high reward. So I think reward seeking is sort of the higher concept and reward hacking is sort of the lower concept.
verbatim transcript · starts at 32:46
32:46Can you can you just distinguish those things? What hacking is when you find unintended tricks or hacks or unintended solutions to problems that essentially developers didn't intend. Um so like a classic example is like in the coastal runners or coast runners I think the game is where it's the the boat is optimiz like the RL agent is optimized to drive a boat in c in in a in a race
33:07and essentially what it learns is just to drift in a corner over and over and maximize reward that way. So that's like one example of like very simple reward hacking. Um now reward seeking is the general um is generally the situational aware reasoning about what is being graded, what is being rewarded and then optimizing your actions for that. So reward hacking is one useful strategy um for getting a high reward. So I think
33:30reward seeking is sort of the higher concept and reward hacking is sort of the lower concept. >> Situational awareness is an important concept. Yeah, situational awareness is important here because I wouldn't say that the boat in the coastal runners has a model necessarily of the environment and what's rewarded. It has like some simple huristics that it's um utilizing to uh get reward. But the difference in the models is they have like a very very
33:52complex model of like what's environment, what are the rubrics, what might they check, what are the edge cases and so on. And that's when it really starts to turn over into reward seeking. I think it's important to say that you you can be reward seeking without being reward hacking and you can be reward hacking without being rewarding or you can be both or you can be neither. So this spinning boat
34:14example is reward hacking. The policy that was learned by the optimizer um is exploiting a loophole in the optimizer but the policy itself does not have an internal representation of this exploit uh or of the fact that it is doing this. Um, a reward seeker uh could also realize that there are very good graders that block all reward hacks and therefore it will never reward hack because that would not actually lead to