Inference has the fastest research-to-production timeline of any technical field, with new techniques going from paper to production support within hours.
Philip Kiely argues that inference engineering moves from research to production faster than any other field, sometimes in hours, citing an engineer who implemented a new quantization paper as a CUDA kernel just 31 hours after it was published. ✦ AI generated
Philip Kiely · The TWIML AI Podcast · 2026-04-30 · original ↗
starts at this moment · 10:31
“Do you find that that research to production timeline in inference is you know, particularly rapid relative to other aspects of AI?”
Even within AI, what moves faster? Training ones, for example, if you want to train a model off of a new technique, it can still take weeks or months to fine-tune the hyper parameters and find the exact right way to sort of express that technique. But with inference, the timeline is often hours. A new model architecture comes out, you have to figure out how to support it day zero.
verbatim transcript · starts at 10:31
10:17about it all over in 2 weeks and everybody's already using it. Like do you find that that research to production timeline in inferences you know, particularly rapid relative to other aspects of AI? I mean, it might be the fastest timeline in the world. If you think about medicine, for example, it can take decades for research to reach a pharmacy. Um if you think about, you know, physics or engineering, it can
10:46take years to apply a new concept, material science. Even within AI, what moves faster? Training ones, for example, if you want to train a model off of a new technique, it can still take weeks or months to fine-tune the hyper parameters and and find the exact right way to sort of express that technique. But with inference, the timeline is often hours. A new model architecture comes out, you have to figure out how to
11:13support it day zero. We had Polo Quant come out, that research paper, and an engineer on our model performance team had it implemented 31 hours later as a as a CUDA kernel. So, the the pace of applied research is very very fast because it is a highly competitive industry and everyone is sort of searching for that next edge. Maybe the only industry where research goes into production faster than
11:42inference might still be like trading. To to be clear, trading with a with a D, not training. Yeah, this might be a little bit too in the weeds, but with regards to Polo Quant, like the paper was a year old and then all of a sudden it got really popular last week or a couple weeks ago. Like what was what's the backstory there? Yeah, so it was sort of republished as a as a blog
12:06post um and sort of caught everyone's attention. I think like as as a little detour, you know, I I grew up in Iowa. So, I was very influenced by sort of the best Midwestern school of thought in economics, which is the Chicago School of Economics. Uh famed creator of the efficient market hypothesis. I think the most dangerous falsehood that you can expose an impressionable young person into is the efficient
- ·Inference moves research to production fastest of any technical field
- ·New model architectures need day-zero support
- ·Training instead takes weeks or months to fine-tune
- ·Engineer implemented new quantization paper as CUDA kernel
- ·Shipped just 31 hours after publication
- ·Contrasts sharply with training's slower iteration cycle