ATRIUMsearch → argument graph
MechanismVideo · 5:00 — 6:30

Cosine can compete with US labs on a fraction of their budget because it licenses model weights for customers to run themselves rather than hosting inference, avoiding the massive data-center spend that consumes most of a frontier lab's capital.

Answering how Cosine can do with millions what US labs do with billions, Alistair says Cosine licenses its technology instead of selling inference tokens, so it avoids the huge data-center costs that dominate spending at labs like Anthropic. ✦ AI generated

Alistair Pullen · Machine Learning Street Talk · 2026-07-13 · original ↗

starts at this moment · 5:00

Elicited by

How can you do in millions what they are doing with billions?

what that means in practice for cosign is that we license the technology that we build. We don't actually make a margin on tokens or anything like that. Um and all of this ties into your question meaning um we don't have to spend a lot of the money that the Americans are having to spend on data centers for inference purposes.

verbatim transcript · starts at 5:00

Transcript · around this moment

5:00model weights um that we give to them deploying them on their own GPUs. We have a lot of that that is like the most air gap the most um secure deployment we do uh or they are renting GPUs in some hyperscala cloud that they're already a part of. So whether it be like Azure or um AWS or whatever and then they'll run the model there. Um what that means in

5:21practice for cosign is that we license the technology that we build. We don't actually make a margin on tokens or anything like that. Um and all of this ties into your question meaning um we don't have to spend a lot of the money that the Americans are having to spend on data centers for inference purposes. Now that's not to say that you don't also need a huge amount of compute for

5:44training. Obviously you do and a huge amount of the infrastructure they have in the US will also be used for training. But I I think that one of the biggest reasons you've seen people like Anthropic struggle recently uh and the reason they've signed the deals they have with like the Colossus cluster and so on is because of inference and not because of training. So you do need

6:03significantly less resource if you're not going to do the inference bit. And we're fortunate in that the way that we sell the product it means we don't really have to. And then on the other side of that um we are taking some interesting research approaches in terms of like how you pull something like this off and we can talk more about that in a minute I'm sure in in terms of how we

6:19are architecting the model, how we're training it, some algorithmic stuff. All of that's to say that we we do have um a credible shot at at pulling off the full run uh including like the continued pre-training, the mid training, the post- training, all of those bits. But to be completely transparent with you, there isn't a huge amount of room for error or a huge amount of wiggle room.

6:40There are obvious places where we have had to make trade-off decisions. You know, that includes that that extends like the scope of RL. I'd always like to do larger RL runs, more generations, um more um you know, inference time compute um during the RL process to get more variety. we can't do as much of that as we would like to if we had like 10 times more compute for instance right

Around this claim