ATRIUMsearch → argument graph
MechanismVideo · 4:34 — 6:04

Cosine can compete with vastly better-funded US labs because it licenses model weights for customers to run on their own infrastructure rather than hosting inference itself, so it avoids the massive data-center inference spend that burdens companies like Anthropic.

Pullen explains that Cosine avoids the huge inference/data-center spend dominating US labs' budgets because it licenses model weights for customers to deploy themselves rather than serving tokens through its own hosted API. ✦ AI generated

Alistair Pullen · Machine Learning Street Talk · 2026-07-13 · original ↗

starts at this moment · 4:34

Elicited by

How can you do in millions what they are doing with billions?

Most of the time now nearly all of the time we are not hosting the model ourselves. A customer isn't hitting you know cosign slash you know APIVv1 and then hitting a chat completions endpoint or something like that from us. They are either taking the model weights um that we give to them deploying them on their own GPUs.

verbatim transcript · starts at 4:34

Transcript · around this moment

4:16you know, let's say 14 >> single to double digit billions. Yeah. >> Something like that. How can you do in millions what they are doing with billions? >> Yeah. Yeah. No, it's it's a very fair question and one that I probably get more than anything else. We at Coine are not an inf an inference company. Uh and that sort of ties into the kinds of deployments and the way that we sell our

4:34product. Um so for your viewers I should probably give a bit of background given the fact that we predominantly deploy into highly secure high-sided environments. Most of the time now nearly all of the time we are not hosting the model ourselves. A customer isn't hitting you know cosign slash you know APIVv1 and then hitting a chat completions endpoint or something like that from us. They are either taking the

5:00model weights um that we give to them deploying them on their own GPUs. We have a lot of that that is like the most air gap the most um secure deployment we do uh or they are renting GPUs in some hyperscala cloud that they're already a part of. So whether it be like Azure or um AWS or whatever and then they'll run the model there. Um what that means in

5:21practice for cosign is that we license the technology that we build. We don't actually make a margin on tokens or anything like that. Um and all of this ties into your question meaning um we don't have to spend a lot of the money that the Americans are having to spend on data centers for inference purposes. Now that's not to say that you don't also need a huge amount of compute for

5:44training. Obviously you do and a huge amount of the infrastructure they have in the US will also be used for training. But I I think that one of the biggest reasons you've seen people like Anthropic struggle recently uh and the reason they've signed the deals they have with like the Colossus cluster and so on is because of inference and not because of training. So you do need

6:03significantly less resource if you're not going to do the inference bit. And we're fortunate in that the way that we sell the product it means we don't really have to. And then on the other side of that um we are taking some interesting research approaches in terms of like how you pull something like this off and we can talk more about that in a minute I'm sure in in terms of how we

6:19are architecting the model, how we're training it, some algorithmic stuff. All of that's to say that we we do have um a credible shot at at pulling off the full run uh including like the continued pre-training, the mid training, the post- training, all of those bits. But to be completely transparent with you, there isn't a huge amount of room for error or a huge amount of wiggle room.

Around this claim