The safe way to specify an AI's reward function is not to hand-write an outcome like 'end world hunger,' but to empirically estimate a reward function from the steady-state distribution of current human actions and outcomes (via maximum entropy inverse reinforcement learning / active inference), then perturb that distribution slightly and evaluate the consequences.
Beck argues that dangerous AI outcomes stem from naively hand-specified goals, and that a safer approach uses maximum entropy inverse reinforcement learning to derive a reward function from current human behavior, then nudges it incrementally. ✦ AI generated
Jeff Beck · Machine Learning Street Talk · 2026-01-25 · original ↗
starts at this moment · 44:55
“How should they make sense of all of this?”
Are you familiar with maximum entropy inverse reinforcement learning? I like to call it active inference because it's really similar. So there what you're doing is you're basically observing someone's policy and then you're trying to do a maximum entropy model. You're doing maximum model on the reward function itself.
verbatim transcript · starts at 44:55
44:55really similar. Um and so there what you're doing is you're basically observing someone's policy and then you're trying to do a maximum entropy um model. You're doing maximum model on the reward function itself. Um at the end of the day what ends up happening when you do this is this is why it's like basically just like active inference. You get a reward fun. So you have some you know organism or whatever
45:17and you're trying to do this for it and and it it's got some stationary distribution over actions and outcomes right it's inputs and outputs of a stationary distribution that becomes your reward function like not directly there's some math involved but basically your reward function is a function of the steady state distributions over actions and outcomes so we could do this right we could take the current we could
45:34take the current manner in which humans are making decisions and we could write down right what's the stationary what what is the current estimate of the stationary distribution of our actions and outcomes. So this would include things like everyone's getting you know this number of people are going hungry this you know and and you know all the stats that describe like the inputs and outputs to our policy make you know to
45:52our policy decision um and then we could just ask an AI your reward function is the one that results in the same outcome that we currently have right on average and it would execute it and it would and and to the extent that it works right it it it would it would ultimately result in a in an AI algorithm that just sort of is like mimicking human behavior, right? Or
46:15it's at least achieving the same outcome that we were achieving before. Now, here's the safe way to like improve the situation. You don't say end world hunger, right? You perturb that distribution >> over outcomes, right? And just just over outcomes a little bit >> and then you evaluate the consequences, right? It's it's all you're doing. You make these little changes in the reward in an empirically estimated reward
46:41function, right? rather than just sort of specifying one by hand because that's the dangerous thing. >> Jeeoff, thank you so much for joining us today. >> It's my pleasure. >> Amazing.
- ·Naively hand-writing outcomes like 'end world hunger' is unsafe
- ·AI systems with fixed goals can produce catastrophic side effects
- ·The reward function itself is the source of risk
- ·Use maximum entropy inverse reinforcement learning (active inference)
- ·Observe the steady-state distribution of current human actions
- ·Empirically estimate the reward function from that distribution
- ·Perturb the estimated reward function slightly
- ·Evaluate the consequences before deployment
- ·Avoids the leap from current behavior to an extreme goal