Building an effective inference system requires broad expertise across many complex, disparate disciplines, similar to how a mixed martial artist must master multiple distinct fighting styles rather than excelling at just one.
Kiely uses an MMA analogy to explain that inference engineers must be competent across GPU programming, applied research (quantization, speculation, KV cache), and large-scale distributed systems all at once. ✦ AI generated
Philip Kiely · The TWIML AI Podcast · 2026-04-30 · original ↗
starts at this moment · 8:23
“Talk a little bit more deeply about like what makes inference difficult.”
You need to have a wide range of expertise on a lot of very different complicated topics to build a truly effective inference system. There's what happens on the GPU, right? There's understanding CUDA level programming and PyTorch on top of that and then the inference engines on top of that.
verbatim transcript · starts at 8:23
8:23striking of various disciplines, I've done wrestling in various disciplines and I've started to understand how those things blend together. I think inference is very similar. You need to have a wide range of expertise on a lot of very different complicated topics to build a truly effective inference system. There's what happens on the GPU, right? There's understanding CUDA level programming and PyTorch on top of that and then the inference engines on top of
8:50that. There's being able to take and apply the research, the understanding of different quantization techniques, different speculation algorithms, different KV cache reuse mechanisms. How do you parallelize models across GPUs? How do you disaggregate across different types of hardware? So, there's all that applied research. And then on top of that, there all the traditional problems of running very large scale distributed systems. Anything that you're familiar with from
9:16the last 20 years of of cloud architecture as well as understanding how to, you know, take very very large volume workloads that are coming from around the world and and handle them appropriately. So, every piece of back end web infrastructure, every piece of GPU programming, it all comes together in inference in windows and and packages that are very demanding. Like you often have a couple hundred milliseconds
9:46latency SLAs that you're dealing with. So, it's the it's the orchestration of a a large number of very complex systems together that makes it such a difficult problem. In AI broadly, like there's this very tight relationship between research and implementation. Uh but in inference in particular, like I I've noted on several occasions, like you know, hear about the speculative decoding work one day and like I hear
10:17about it all over in 2 weeks and everybody's already using it. Like do you find that that research to production timeline in inferences you know, particularly rapid relative to other aspects of AI? I mean, it might be the fastest timeline in the world. If you think about medicine, for example, it can take decades for research to reach a pharmacy. Um if you think about, you know, physics or engineering, it can