ATRIUMsearch → argument graph
ClaimVideo · 30:48 — 32:18

Running AI models locally doesn't actually save money, because any efficiency breakthrough that makes local hosting cheaper makes cloud hosting even cheaper; the real benefit of local models is privacy, not cost.

Dax argues local-model cost savings are mostly illusory since any hardware/model efficiency gain benefits cloud providers even more, making local inference primarily a privacy play rather than a cost-saving one. ✦ AI generated

Dax Raad · Syntax · 2026-07-15 · original ↗

starts at this moment · 30:48

Elicited by

And what about the crazy people that think you can run it locally? Um including us, what's your take on the people that think they're going to have a machine running it in their in their backyard?

if you focus on cost, it's not really like the local model isn't going to help you on cost. Cuz any mechanism that makes it cheaper to host locally makes it like 10x cheaper to host in the cloud. Like if a model gets more efficient or more capable at a smaller size, that's just going to be cheaper per token in the cloud. So, I think local model is more of a privacy thing, less so a cost thing.

verbatim transcript · starts at 30:48

Transcript · around this moment

30:48just do not want stuff outside of your house, makes total sense. You can kind of go down this path. Um but if you think about the underlying but if you focus on cost, it's not really like the local model isn't going to help you on cost. Cuz any mechanism that makes it cheaper to host locally makes it like 10x cheaper to host in the cloud. Like if a model gets more

31:09efficient or more capable at a smaller size, that's just going to be cheaper per token in the cloud. So, I think local model is more of a privacy thing, less so a cost thing. And just to kind of give you guys a few numbers. So, like again because we are an inference provider, we see some of these details. Um we still use middlemen. Despite using middlemen for hosting our

31:32GPUs, there are some models that we are able to host at a 70% discount to us. Um that is very very cheap, which means we can make 70% margin by selling it at sticker price. And that's with a middleman involved. So, if you directly spend the capital to acquire the GPUs, uh you can probably hit those like 90% margins that I'm estimating for Anthropic. Which means like the cost of inference

31:57is is very very cheap. Yeah. But again, these are for open-source models. >> Yeah. >> So, we have we still have to rely on those getting better. Um but you know, it's going in that direction so far. >> Yeah, that's encouraging. Uh let's talk about Claude Code. So, Claude Code it seems like it's been very unclear of what their stance is, um whether or not even providers like Open Code can use

32:23things like the Claude Code Max plan. Or it just seems like they are probably a difficult partner to work with. But what is the current status of Claude Code, the Claude Code Max plan, Open Code, and third-party harnesses. >> Yeah, so the integration that the plugin that we had in open code that let you use your max plan, that's definitely not allowed. They we fought with them on that for a lot

32:48and we did not win. So, [laughter] that that that's definitely not allowed. Of course, people still find ways to hack it in. We just can't officially support it. Um the the SDK, which is like spawning Claude or like using Claude headlessly, that is now in a gray area. They're saying for now it's allowed. So, a product conductor can wrap it. A product like T3 code can wrap it.

Around this claim