ATRIUMsearch → argument graph
MechanismAudio · 17:12 — 20:00

Enterprise agents can achieve verifiability by leveraging the system of record (the database) to define expected outcomes, but to progress toward autonomous agents, you must capture the 'tribal knowledge' that lives in people's heads or Slack channels — creating a data flywheel where every agent interaction generates new data for evals and process improvement.

Philipp describes a two-lane approach: first, use the existing system of record (the database) to verify agent outcomes by checking expected results. Second, agents ask users clarifying questions, and those decision traces get stored — turning 'process mining' into 'agent mining.' This creates a flywheel where captured data becomes new evals, which can either flag anomalies or be elevated into new standard operating procedures. ✦ AI generated

Philipp Herzig · No Priors · 2026-04-23 · original ↗

plays this moment only · 17:12 — 20:00

Elicited by

Do you think it's possible in terms of verifiability or the ability to go understand and evaluate against that intent? Because it is much, I don't know if I would say it's more diverse than code, but it's not obviously verifiable, as you pointed out. Like, do you think it can be?

That's exactly the point. That is where the starting condition is great, right? I think in terms of like 2 lanes. The first lane is, of course, you have the system of record today, right? You know exactly. in the system, hey, given this or that instruction, right, what is the outcome, right? Because you can see it in the database, right? And then you can construct, hey, if the order to cash process runs like this, then you need to expect, right, that the cash, like the accounts receivable, needs to come in this way, right, with the following taxes and so on and so forth. So that gives you verifiability. Now, the challenge, of course, is, rather, this is never enough, right? Because If you just look into the system of record today, that data is insufficient for this grand vision that everybody has, that it becomes this autonomous enterprise, or the agency of these agents is increasing over time. So at the beginning, the agents, of course, are coming back to you. Some people call this human in the loop or whatever, right? So they need to come back to you, like also still with Cloud Code or Codex, and still ask you some clarifying questions. Hey, I could now go this way, I could do that way. And with that, what you want to design for is that you start to capture more of that context. I always call this the tribal knowledge, the stuff that is not in the system, stored somewhere, that just lives in people's heads or maybe in Slack channels, maybe in Teams channels, maybe it was just a discussion on the phone. So it's not stored anywhere. So how can you drive a decision from that? And then so the question is, the agent needs to come back, ask you for input, now you wanna store that. And now what we do, in the past we called this process mining, now we call it agent mining, because you record all these decision traces, these contexts, what the users are entering into the system. And then you can either use it to say like, Hey, wait a minute, this is actually an anomaly. The folks in, I don't know, in UK from our company or the folks in Australia shouldn't do this because the standard operating procedure is this. Or you say like, that's actually a very good improvement. And then you can elevate this to be the new standard operating procedure, maybe not just for Australia, but maybe for the rest of the world or more countries to run your company more efficiently, because now you'll learn something how the organization behaves, because it can go two ways, right? It could be either good. Or it could be a bad thing, and then you maybe want to streamline the process, how people then actually conduct the process in a different way. And that then leads to this kind of, I call this then this data flywheel, so to speak.

verbatim transcript · starts at 17:12

Around this claim
Extends · 4
This moment responds to