ATRIUMsearch → argument graph
DefinitionVideo · 8:28 — 9:51

VLM is an inference engine whose job is to turn available GPUs into a running endpoint for intelligence, making it critical software — like databases and operating systems — that powers how people build and run AI models.

Simon introduces what VLM is and where it sits in the AI stack: it is an inference engine that turns any available GPU into a running endpoint for intelligence, analogous to databases and operating systems as critical software powering the AI economy. ✦ AI generated

Simon Mo · a16z Podcast · 2026-08-06 · original ↗

starts at this moment · 8:28

Elicited by

And can you talk about where VLM sits in that stack where where we do have these larger enterprise companies that are choosing to use open source models like where where does VLM sit in the stack for them?

VLM is a inference engine. That means its job is to turn available GPUs into a running endpoint for intelligence. So that means it is kind of like databases and operating system other critical software to power uh this uh economy or power of the AGI that everybody really uses today to ensure they can have uh cost effectiveness, efficiency, reliability and also always staying on the frontier because for VRM we support more than a thousand model architecture up to today and a lot of those are proprietary But also a lot of those are openweight right...

verbatim transcript · starts at 8:28

Transcript · around this moment

8:08these products but you know some of the most innovative products and applications now um you know really depend on this very deeply. >> Yeah. Yeah. And can you talk about where VLM sits in that stack where where we do have these larger enterprise companies that are choosing to use open source models like where where does VLM sit in the stack for them? >> Yeah. I mean you should like just about

8:28everybody uses VLM. You should describe it. >> Just about everybody uses VLM. VLM is a inference engine. That means its job is to turn available GPUs into a running endpoint for intelligence. So that means it is kind of like databases and operating system other critical software to power uh this uh economy or power of the AGI that everybody really uses today to ensure they can have uh cost effectiveness,

8:59efficiency, reliability and also always staying on the frontier because for VRM we support more than a thousand model architecture up to today and a lot of those are proprietary But also a lot of those are openweight right and a lot of those model architecture when they're becoming transitioning from a research prototype to world accessible open way model architecture they are live on VM on immediately so that's what a process

9:27we call day zero model release and additionally VM also work closely with all the hardware vendors so that means across like Nvidia AMD Google and Amazon Intel and lot more their newest chip will make sure VM can run on them and then a lot of cases they use VM as a benchmark to make sure it runs well on them. So this kind of fusion of where models run and where hardware where it

9:51gets gets to meet the hardware is where the magic happen and this is where VM is >> and you've told me some of the behind the scenes stories like it's actually not easy these days model releases it's like a lot of human drama in addition to like technical work I guess are there any stories there that you think are okay to share? Oh, it's actually a very

10:08fun co-design process because from model labs point of view, right, these are brilliant researchers who have built this model. Now their biggest question becomes how do we get this out of the world and make sure everybody's able to use it and run it well. And we have worked with model labs that are um very just because they just use VM already in production or in their research process.

Related moments