ATRIUMsearch → argument graph
FactVideo · 28:17 — 29:47

Older GPU generations like Hopper retain strong demand and rising rental prices years after release, due to export-control-driven Chinese open-source optimization for Hopper, its FP8 quantization support, and its efficiency for smaller models.

Kiely explains that Hopper GPUs remain highly sought after even as Blackwell rolls out, partly because export controls push Chinese labs to optimize for Hopper, and because Hopper (unlike Ampere) supports FP8 quantization cost-effectively. ✦ AI generated

Philip Kiely · The TWIML AI Podcast · 2026-04-30 · original ↗

starts at this moment · 28:17

Elicited by

I'm just wondering if you have like a strong feeling about the GPU lifespan as it relates to inference and inference maturity and the way the field is going.

The thing is the time it takes for a GPU generation to first be manufactured and actually distributed to the point where everyone can get their hands on it is months to a year. And then there's another cycle of months to a year of porting all of the code in the industry to run on it. So actually Hopper GPUs in particular still are very very popular for inference.

verbatim transcript · starts at 28:17

Transcript · around this moment

28:17generation to first be manufactured and actually distributed to the point where everyone can get their hands on it is is months to a year. And then there's another cycle of months to a year of porting all of the code in the industry to run on it. So actually like Hopper GPUs in particular still are very very popular for inference. One big reason for that actually is because so

28:40much open source work comes out of Chinese labs who due to export controls generally work on Hopper GPUs and not Blackwell GPUs. So you generally get like either, you know, FP8 or int 4 kernels, you get uh things that are built for Hopper's asynchronous programming paradigm instead of the slightly different paradigm of of Blackwell kernels. You get things that are models that are built for the size and restriction of say like an

29:078x H200 node instead of say a a GB300 NVL72 system. So there's a lot that still is built for Blackwell. Uh sorry for Hopper. I think that like Ampere has finally fallen out of favor in inference for the most part, mostly because it doesn't have support for FP8 quantization, so you have to run models at full precision or suffer sort of catastrophic quality loss. And that makes Hopper GPUs like much

29:40more cost effective for inference versus versus um Ampere, but due to both like Blackwell shortages as well as the large number of sort of Hopper optimized workloads in the industry, like Hopper has had a ton of staying power. It you know, an H100 is more expensive now than it was a year ago uh in terms of a a rental basis. And I think that like as we continue to

30:09radically underestimate the demand for inference, um even though we're able to, you know, scale models bigger, make them cheaper, make them faster, there's going to be a very strong sustained demand for Hopper inference. I think there's going to be a very very strong sustained demand for Blackwell inference even as Reuben rolls out. And yeah, the other thing about Hopper is that they are really good for like

Related moments