Cosine can compete with labs spending billions because it licenses model weights to customers rather than hosting inference itself, avoiding the massive inference-serving costs that are straining even Anthropic.
Alistair argues Cosine sidesteps the huge inference infrastructure costs that plague US labs because its business model licenses model weights for customers to run themselves, rather than serving tokens at scale. ✦ AI generated
Alistair Pullen · Machine Learning Street Talk · 2026-07-13 · original ↗
starts at this moment · 5:21
“How can you do in millions what they are doing with billions?”
what that means in practice for cosign is that we license the technology that we build. We don't actually make a margin on tokens or anything like that... I think that one of the biggest reasons you've seen people like Anthropic struggle recently uh and the reason they've signed the deals they have with like the Colossus cluster and so on is because of inference and not because of training.
verbatim transcript · starts at 5:21
5:21practice for cosign is that we license the technology that we build. We don't actually make a margin on tokens or anything like that. Um and all of this ties into your question meaning um we don't have to spend a lot of the money that the Americans are having to spend on data centers for inference purposes. Now that's not to say that you don't also need a huge amount of compute for
5:44training. Obviously you do and a huge amount of the infrastructure they have in the US will also be used for training. But I I think that one of the biggest reasons you've seen people like Anthropic struggle recently uh and the reason they've signed the deals they have with like the Colossus cluster and so on is because of inference and not because of training. So you do need
6:03significantly less resource if you're not going to do the inference bit. And we're fortunate in that the way that we sell the product it means we don't really have to. And then on the other side of that um we are taking some interesting research approaches in terms of like how you pull something like this off and we can talk more about that in a minute I'm sure in in terms of how we
6:19are architecting the model, how we're training it, some algorithmic stuff. All of that's to say that we we do have um a credible shot at at pulling off the full run uh including like the continued pre-training, the mid training, the post- training, all of those bits. But to be completely transparent with you, there isn't a huge amount of room for error or a huge amount of wiggle room.
6:40There are obvious places where we have had to make trade-off decisions. You know, that includes that that extends like the scope of RL. I'd always like to do larger RL runs, more generations, um more um you know, inference time compute um during the RL process to get more variety. we can't do as much of that as we would like to if we had like 10 times more compute for instance right
7:05so there are trade-offs but fundamentally I think given the way that we've scoped the project um and we can get into more detail I'm sure but given the way that we've scoped the project I think it is it is viable in that in that very narrow scope we have some of the largest companies in the UK all feeding use cases and you know their desires for what they want the model to be able to