ATRIUMsearch → argument graph
MechanismVideo · 25:52 — 27:22

Effective evaluation for multi-agent systems must be end-to-end rather than per-agent, because evaluating an individual agent in isolation is meaningless unless it also works correctly upstream and downstream within the whole system.

Rashmi argues that eval frameworks for agentic systems must shift from evaluating individual agents to evaluating the whole pipeline end-to-end, since isolated agent performance doesn't guarantee system success. ✦ AI generated

Rashmi Shetty · The TWIML AI Podcast · 2026-04-16 · original ↗

starts at this moment · 25:52

Elicited by

And then you mentioned uh evals earlier. Talk a little bit about how the teams there approach evals for these types of systems.

Eval frameworks are basically tuned to do end-to-end evals rather than individual agent evals, because individual agent evals gives you nothing unless it works upstream and downstream for the whole system.

verbatim transcript · starts at 25:52

Transcript · around this moment

25:52eval systems. But, the fundamental principle is very much the same, right? You have your golden data sets, you have your specific matrices that you want to kind of um adhere to or or hit, and your experimentation evolves around that. The only difference is that now eval frameworks are basically tuned to do end-to-end evals rather than individual agent evals, because individual agent evals gives you nothing unless it works

26:17upstream and downstream for the whole system. So, that's the that's the that's the nuance difference um that that that So, you have to kind of come up with your own evaluation framework for your for your um for your use case. But, as a platform, we offer different eval uh tools that is needed for you to a build your eval pipelines, um deploy them, provision that uh golden data set,

26:47offer you offer those sandboxed environments where this can be executed. How do you approach the speed at which models evolve in the space? Yes, this is not unique to Capital One, right? But, there is a there is a there is a there's a nuance nuance there as well in terms of Capital One's platform risk-first platform strategy coming to help us benefit us in this specific space, right? As

27:16uh as uh new models come into the fore, uh there are specific layers that has been built in in terms of, you know, provisioning inference layers, inference optimization layers, and benchmarkings uh that makes the evaluation of new models quicker and easier easily provisioned within the platform infrastructure. So, whenever there's a need and there's a new kid in the block, we have mechanisms to uh rapidly experiment on it, deploy and experiment

27:48on it. And does the presence of inference and inference optimization imply that you tend to favor uh self-hosted models? This is another core underlying philosophy of Capital One, right? When you uh you are the most successful when you can offer two things, reasoning and specialization. So, reasoning capabilities with our agentic platforms agentic frameworks in the platform, we are bringing that to the fore. Specialization is something that is very very crucial, right? So,

Around this claim