ATRIUMsearch → argument graph
Video · 2026-07-25 · 1h 3m · 6 moments

Mercor CPO on Revenue Concentration from Frontier Labs

✦ AI generated

timeline · colored by role

01
Claim

Open source models just raise the floor of what people are interested in; data is most valuable at the frontier of model performance, not for capabilities that open models already handle.

Oswald argues that open source models don't threaten Mercor's business because customers need data to push the frontier of model capabilities, not to replicate what open models can already do.

transcript

Oswald Nitski: I wouldn't say that open source model improvements cannibalize our core business because data is most valuable on the frontier of model performance. So each of our customers has their own unique goals and is purchasing eval training data sets to fill gaps in current model capabilities. Open models just raise the floor of what people are interested in. As long as customers still have new capabilities that they want to get better at, our business still continues to grow. Open source models just mean that nobody's buying anything that can already be done.

02
Claim

The percentage framing of what models can and cannot do is totally off; we need to think about continuous uncapped rewards for workflows like legal arguments and medical advice where you could always get better.

Oswald argues that binary percentage-based thinking about AI capability misses the point — for domains like legal and medical reasoning, there is no ceiling, and the right framework is continuous improvement.

transcript

Oswald Nitski: Well, in our Apex benchmarks, we're getting closer to around 50% of long horizon workflows. Top models are scoring around that much. But I think that there's a class of workflows that are just sufficiency-based where you do it and it's done and you're good. This is something like updating a CRM. You couldn't really get much better at it. And then there's a class of workflows that we shouldn't even be thinking about in terms of binary, like, can the models do it or not. And these can be things like legal arguments or to an extent medical advice, where you could always get better. And in those cases, I think that the percentage framing is just totally off and we need to be thinking more about continuous uncapped rewards.

explains mechanism · 1extends · 1gives example · 1

03
Claim

There is no enterprise ROI problem with AI right now; we are in a period of exploration and experimentation with patience for ROI calculations to settle.

Oswald argues that enterprise AI spend faces no ROI crisis — companies are still in an exploration phase, tolerating uncertainty because the landscape is shifting too fast for ROI calculations to stabilize.

transcript

Oswald Nitski: I don't think there's an ROI problem right now. I think we're in a period of exploration and experimentation where there's more tolerance, more patience to get that ROI calculation right now. There's a lot of different projections around where token prices will go, where performance will go, and right now we're starting to see some amount of tightening of the screws on spend here and there. But I think the paradigm we're in is still let's see what happens because things are moving so quickly that the ROI calculation might shift too dramatically still.

04
Prediction

Mercor's biggest challenge is moving down-market to diversify revenue away from frontier model labs, because serving smaller enterprises with efficient self-serve human data projects is a harder product to build.

Oswald acknowledges that Mercor's revenue is concentrated among frontier labs, and the company's strategic priority is building self-serve products to serve the broader enterprise market.

transcript

Oswald Nitski: So, I can answer this from a how it affects the product team. We would love to move like our biggest challenge is moving down market so that every single enterprise can efficiently run human data projects for eval and training. And that'll diversify our revenue for sure because there's many more enterprises than there are labs. And that's a harder product to build. And that's the direction that we are taking our products, taking the company — to be able to self-serve projects very efficiently, have like AI project managers so that it's a lot easier to do this work for smaller customers because running a human data project for a lab is incredibly hard. It's a white glove service that requires a lot of people on the operations team. As we make that more efficient with better products, better processes, we can do smaller projects that are more heterogeneous for more customers. It's the direction we have been heading which has reduced concentration and it's the direction that we'll continue to head as every enterprise begins to have human data work for their proprietary use cases.

05
Prediction

RL environments — simulations of apps that agents need to use — are the fastest-growing data type and represent the next frontier for model evaluation and training.

Oswald identifies 'environments' — simulated worlds that train agents to use tools like Salesforce — as the fastest-growing data category, representing a shift toward deployment-relevant evaluation data.

transcript

Oswald Nitski: The data type that's growing the fastest for us is environments. People you know that you might have seen a lot about these RL environments on Twitter. It's kind of like a hype term. Every company kind of has a different definition for it. But we are certainly the leader in the category and view it as basically these simulations of apps that you might want your agent to use. And also as per rich start state, which we call like the world that is basically representative of all the data you might have on your machine like your laptop. And then we have tasks that train agents how to use those tools to accomplish something that's useful. It's a bit of a complicated annotation process because the agent has to interact with this simulated world. We have to make that start state, which can be hundreds of files, thousands of files. And the shift here is that the data that the models are now the agents are being evaled and trained on looks a lot closer to what they see in deployment. So, if you want to learn how to use something like Salesforce, you need a pretty high-fidelity mock that acts exactly like Salesforce in your eval and training. And it's complicated to get this set up. Just like years ago preference ranking was really hard to get set up. SFT was really hard to get set up when InstructGPT first came out. So, this is the frontier right now. Labs are figuring it out. Eventually, it'll get so smooth that enterprises can do it, too.

gives example · 1provides context · 1

06
Prediction

Evals and training data are the primary bottlenecks in model performance, and as long as better models are valuable to the economy, demand for them will grow — making Mercor a tech-enabled services company worth $200 billion.

Oswald argues that since eval and training data are the binding constraint on model performance, and demand for better models will only increase, Mercor's core business has massive growth potential.

transcript

Oswald Nitski: We're basically like a tech-enabled services company. Our services are incredibly valuable in driving revenue gains for our customers primarily through better model capabilities. Evals and training data are the primary bottlenecks in model performance right now. If every enterprise needs to have specialized proprietary models, even if the capabilities start to saturate, the eval serves as the PRD for exactly what you want, but also the optimization objective for better performance. As long as better models are valuable to the economy, there will be demand for eval sets and training sets. If we can make that process faster and faster, we can serve a growing demand for human data eval training. And then we also have a growing agent deployment enterprise arm as well.

explains mechanism · 1gives example · 1supports · 2

Highlight slides
Related episodes