ATRIUMsearch → argument graph
MechanismVideo · 2:37 — 4:07

Optimizing inference requires a cost-per-token and architectural strategy, combining memory technologies like KV cache and CXL with intelligent routing tools like NVIDIA's Dynamo to balance workloads across GPU, CPU, or LPU types.

Kaushik Shirhatti explains that inference optimization is a cost-and-architecture problem solved with memory technologies and intelligent routing across mixed compute types. ✦ AI generated

Kaushik Shirhatti · The TWIML AI Podcast · 2026-07-02 · original ↗

starts at this moment · 2:37

Elicited by

talk a little bit about what you're seeing on the inference side among your customers.

You have to look at inference from a cost optimization perspective, power to token, and then also from a, let's say, architectural strategy perspective. So you can see the adoption of inference with technologies like memory optimizations, like KB Cache, CXL, intelligent routing capabilities like Dynamo from NVIDIA as well

verbatim transcript · starts at 2:37

Transcript · around this moment

2:37on the inference side among your customers. You have to look at inference from a cost optimization perspective, power to token, and then also from a, let's say, architectural strategy perspective. So you can see the adoption of inference with technologies like memory optimizations, like KB Cache, CXL, intelligent routing capabilities like Dynamo from NVIDIA as well And so the ability to optimize, from the request all the way to the

3:06token, is load balancing cost fit for purpose for that workload in terms of the type of GPU, the type of CPU, or the type of LPU you're using in combination, and the type of memory strategies. it's one thing to use AI via an external service provider. It's a totally different, ball of wax to build out that infrastructure within your organization. What challenges are you seeing for the organizations that take this on?

3:35I think the challenges are, so uniform, so it starts with data as an example, right? So AI is really not meaningful without data that is good data, right? garbage in, garbage out type of thing. So- That's your context. Exactly. So I think that's still a huge challenge for organizations is, one is mapping the data, just knowing where the data is, what format is it in, what sequence is

3:57it in, what do I need to do to curate it, cleanse it, prepare it, and all that. That's still a huge, huge challenge, right? And then protecting the data as well, as you mentioned. but then also I think it's, We're seeing a, an uplift, I think, in ROI, right? So I think a year ago, NVIDIA has done some independent studies and others as well, and we were at the five percent range for ROI.

4:17I think we're now moving up to the 30% range, depending on which segment you're looking at. But as a generality, we're seeing a, big uptick, I think, in, in true ROI return, implementing AI into the organizations. So that process of understanding where to apply AI, I think, has shifted from experimentation and for the sake of AI to now let's look at problem statements and figure out how does AI agentic systems solve

Around this claim