An AI factory, regardless of scale, is a system where you put data and energy in and get intelligence, tokens, and business outcomes out, organized across five layers: energy, chips, infrastructure, models, and applications.
Kaushik Shirhatti defines an AI factory not by size but as a five-layer system spanning energy, chips, infrastructure, models, and applications that turns data and power into business outcomes. ✦ AI generated
Kaushik Shirhatti · The TWIML AI Podcast · 2026-07-02 · original ↗
starts at this moment · 0:46
“When you talk about an AI factory at scale, like what are we actually talking about?”
when you put data in, you put energy in, and boom, like magic, intelligence, tokens, business outcomes come out. The way HPE and NVIDIA thinks about this AI factory is essentially across five layers. So you think about the energy You think about the chips, so all the innovation that NVIDIA brings in, all the innovation that HPE brings in.
verbatim transcript · starts at 0:46
0:46matter if it's big or small, is when you put data in, you put energy in, and boom, like magic, intelligence, tokens, business outcomes come out. The way HPE and NVIDIA thinks about this AI factory is essentially across five layers. So you think about the energy You think about the chips, so all the innovation that NVIDIA brings in, all the innovation that HPE brings in. You think about the infrastructure,
1:16w- you know, all the hardware, all the-- where all the different options we give customers to go hosted. And then you start thinking about different models, and ultimately, the applications. Thierry, how does that jive with the way you think about it? HPE's historically had a private cloud AI offering. You've got a AI factory offering. Are those the same? Are they different? If you think about the scale, the
1:41maturity of different customers, we have customers that need sixteen GPUs, and they need to do generative AI, and they're just getting onto their journey. We have other customers that have been doing HPC for twenty, thirty years and understand traditional HPC environments. Now they're moving into AI and using AI to transform HPC, but they're highly sophisticated. So the concept of AI factory, can move along this medium, in
2:03terms of need and sophistication. if we look back a couple of years, training huge models was the big driver, right? the frontier labs, would largely work with the hyperscalers and contract, many s- servers and GPUs to do training workloads. More recently, though, inferences kinda come to the forefront as the driver for, enterprise and really, creating these new opportunities for neo clouds and others. talk a little bit about what you're seeing
2:37on the inference side among your customers. You have to look at inference from a cost optimization perspective, power to token, and then also from a, let's say, architectural strategy perspective. So you can see the adoption of inference with technologies like memory optimizations, like KB Cache, CXL, intelligent routing capabilities like Dynamo from NVIDIA as well And so the ability to optimize, from the request all the way to the
- ·Not defined by scale, but by function
- ·Inputs: data and energy
- ·Outputs: intelligence, tokens, business outcomes
- ·Energy powers the entire system
- ·Chips: NVIDIA's compute innovation
- ·Infrastructure, models, and applications layer on top