ATRIUMsearch → argument graph
ClaimVideo · 1:41 — 3:11

The real bottleneck in production AI isn't performance or benchmark overfitting, it's reliability, trustworthiness, and understanding what your agents are doing.

Scott Clark argues that the core problem enterprises face with AI is not squeezing extra performance from models, but ensuring they behave reliably and understandably in production. ✦ AI generated

Scott Clark · The TWIML AI Podcast · 2026-05-07 · original ↗

starts at this moment · 1:41

Elicited by

What have you been up to? What's Distributional up to?

one of the things that's really holding back value in the enterprise, especially when it comes to AI, isn't necessarily just performance. People don't stay up at night cuz they're trying to over fit it in eval by another half a percent... what you really care about is is this model or agent going to perform well in production? Is it going to treat my customers the way I want them to be treated and really represent the business well?

verbatim transcript · starts at 1:41

Transcript · around this moment

1:47I came to the realization that one of the things that's really holding back value in the enterprise, especially when it comes to AI, isn't necessarily just performance. People don't stay up at night cuz they're trying to over fit it in eval by another half a percent. Um you can bench max as much as you want, but what you really care about is is this model or agent going to perform

2:08well in production? Is it going to treat my customers the way I want them to be treated and really represent the business well? And that's less about over fitting a benchmark and that's more about reliability, trustworthiness, and really understanding of what your agents are doing. And so, when I left Intel um after a few years there and started Distributional, the original concept of the company was how do we build better tests that are

2:35statistical and Bayesian distributional tests for these AI systems to really take into account all of the stochasticity and chaos and non-stationarity that these models are injecting into these systems in a way that wasn't necessarily true with traditional neural networks or gradient boosted decision trees. It turns out that that was actually not necessarily the real bottleneck. People weren't afraid to deploy these systems. They wanted to learn online. The space was

3:02moving too fast to really be hindered by tests. But, what people rapidly realized was they need to learn very quickly in production. And so, about a year ago we shifted the company from pre-production testing into post-production analytics. And this is really about finding all of the signals and patterns in production data that you might not be catching with a monitoring system today so that you can create these feedback loops to kind

3:30of self-improve and self-heal these agents, helping you find these unknown unknowns through analytics and unsupervised learning and leverage that to make the systems more and more aligned to what you actually care about. >> It's no surprise to hear Bayesian statistics come up in that. That's been a focus of yours going way back. Was it to Yelp? >> Yeah, all the way back to my PhD. >> And talk a little bit about how that

Around this claim
Evidence · 5