Non-stationarity of foundation models is one of the strongest arguments for keeping evaluation in an online loop, because evals that worked yesterday may not work tomorrow.
Clark argues that because underlying models shift over time — even without version changes — any static eval or guardrail will eventually fail, making continuous online analytics essential. ✦ AI generated
Scott Clark · The TWIML AI Podcast · 2026-05-07 · original ↗
starts at this moment · 37:06
“If you could have predicted in a live system that the underlying model distribution is shifting under you — is this kind of analytics part of the answer to that?”
this non-stationarity, which is what you're pointing at here, is actually one of the strongest arguments for keeping this in an online loop. Because whatever evals and guardrails or whatever it was that worked before might not work tomorrow because the model may have shifted around it... It's like you're trying to box something into something in some high-dimensional box and it's going to find some dimension that maybe you didn't realize.
verbatim transcript · starts at 37:06
37:06which is what you're pointing at here, is actually one of the strongest arguments for keeping this in an online loop. Because whatever evals and guardrails or whatever it was that worked before might not work tomorrow because the model may have shifted around it or whatever it may be. It's like you're trying to box something into something in some some high-dimensional box and it's going to it's going to find some dimension that
37:35maybe you didn't realize or it's going to come up with some new skill or it's going to come up with something and they were using reinforcement learning to teach it to like stop talking about goblins, but it ended up learning that it should like be mean to your users or something like that. >> And maybe it's an important point that this is not an anomaly. Like this is by
37:54design. This is what the thing does and why it works so well. >> Yeah, yeah, exactly. And it's it's important to think about these as these dynamic systems that do effectively evolve over time. And yes, Anthropic and OpenAI get to choose the stimulus and the reward that like guides that evolution, but that evolution's going to have some interesting side effects to it. And so I go back to the
38:19there's this fun quote from Jurassic Park from 25 years ago or whatever about like the Raptors like testing the fence and like constantly trying to like find its way out and it's like a fence that works in one situation isn't necessarily going to work in the other or like eventually there's going to be ways to get out of these systems. And this isn't even with like an adversarial system
38:38that's like intentionally trying to escape the box like these things are just going to naturally grow in a way where it's like yeah, it turns out your chicken wire fence doesn't work for an elephant. >> And so like I've often you know when I when I come across these reports like the Anthropic report like I've often wondered where is like the the internet like or the model weather
39:02report like how's my model doing today? Like does this type of approach to analytics give that to me for the types of you know for the models that I'm building on? >> Yeah, well and again I think it's to to use the weather report analogy. I think that there's like two ways to look at this like one is like what what's the temperature today? And that's your