Cosine can compete with vastly better-funded US labs because it licenses model weights for customers to run on their own infrastructure rather than hosting inference itself, so it avoids the massive data-center inference spend that burdens companies like Anthropic.
Pullen explains that Cosine avoids the huge inference/data-center spend dominating US labs' budgets because it licenses model weights for customers to deploy themselves rather than serving tokens through its own hosted API.
transcript
Alistair Pullen: Most of the time now nearly all of the time we are not hosting the model ourselves. A customer isn't hitting you know cosign slash you know APIVv1 and then hitting a chat completions endpoint or something like that from us. They are either taking the model weights um that we give to them deploying them on their own GPUs.
supports · 1