ATRIUMsearch → argument graph
Video · 2026-07-25 · 1h 3m · 6 moments

Mercor Head of Product on Revenue Concentration from Frontier Labs

✦ AI generated

timeline · colored by role

01
Claim

Open source model improvements do not cannibalize our core business because data is most valuable on the frontier of model performance — open models just raise the floor, while customers still need data for new capabilities beyond what open models can do.

Oswald argues that open source models don't hurt Mercor's business because data demand is driven by frontier capabilities, not commoditized tasks. Open models raise the floor, but customers still pay for data to reach new heights.

transcript

Oswald Nitski: I wouldn't say that open source model improvements cannibalize our core business because data is most valuable on the frontier of model performance. So each of our customers has their own unique goals and is purchasing eval training data sets to fill gaps in current model capabilities. Open models just raise the floor of what people are interested in. As long as customers still have new capabilities that they want to get better at, our business still continues to grow. Open source models just mean that nobody's buying anything that can already do what open models can already do.

supports · 1

02
Fact

The 90/10 split between open-source-suitable and frontier-requiring enterprise workflows is wrong — the real opportunity is in latent demand for long-horizon tasks that nobody is even trying to do with models yet, and the frontier models today only handle about 50% of those.

Oswald rejects the common framing that 90% of enterprise workflows are handled by open models. He argues that latent demand for long-horizon tasks (like fully autonomous procurement agents) is unaccounted for, and even frontier models only score ~50% on these complex workflows in Mercor's benchmarks.

transcript

Oswald Nitski: I'm not convinced that 90% of enterprise workflows can be handled by open models or frontier models right now. We think that these calculations might be based off of existing demand or things that come top of mind when current model users are thinking about what models could do, but there's a whole category of latent demand that people aren't even trying to do with models yet. Most commonly we think these are long horizon tasks like setting up a procurement agent to fully automate your procurement team for months on end. You only check on it maybe once a week. We think that's just not even captured in these calculations when someone says enterprise workflows are being handled because nobody's trying to do these things yet. The market for data to support those use cases is growing and that's where we see a lot of the leaders moving to.

extends · 1rebuts · 1

03
Claim

There is no enterprise ROI problem with AI today — we are in a period of exploration and experimentation where there is tolerance and patience to get the ROI calculation right because things are moving too quickly for the calculation to be stable.

Oswald argues the current AI investment climate is not facing an ROI crisis. Instead, enterprises are in an exploration phase, willing to tolerate uncertain returns because the pace of change makes any ROI calculation potentially obsolete quickly.

transcript

Oswald Nitski: I don't think there's an ROI problem right now. I think we're in a period of exploration and experimentation where there's more tolerance, more patience to get that ROI calculation right now. There's a lot of different projections around where token prices will go, where performance will go, and right now we're starting to see some amount of tightening of the screws on spend here and there. But I think the paradigm we're in is still let's see what happens because things are moving so quickly that the ROI calculation might shift too dramatically still.

explains mechanism · 1provides context · 1

04
Mechanism

AI makes the product management job harder, not easier — the challenge shifts from execution to deciding what not to build, as rapid engineering velocity creates pressure to expand surface area chaotically while the real job is fighting to simplify and find the most scalable workflows.

Oswald explains that AI's acceleration of engineering output makes product management more difficult. The bottleneck moves from building features to understanding user needs and exercising judgment about what not to build, as the product team constantly fights against expanding surface area.

transcript

Oswald Nitski: This paradigm makes the job of product management a lot harder because we're trying not to build 10x more product surface area. It makes things incredibly chaotic. We have moments in time where product surface area rapidly expands because people think, 'Oh, I can make all these features really quickly. This is like I could you know, like let me just like push these multi-thousand line PRs.' But we are constantly in this battle to try to simplify our product surface area and find the interactions and the workflows that are most scalable. So, the trend that we see is we're as a product team constantly fighting to reduce surface area and simplify things. And we also see a higher ratio of PMs to eng because engineering is less bottlenecked. There's much, much more work to be done in understanding the workflows of users, the needs of users, and what products actually drive revenue the most becomes the bottleneck now to servicing more demand for us.

05
Mechanism

The secret to scaling supply on the marketplace is a great expert experience — paying experts well and on time, treating their work as dignified, and maintaining strong communication — which drives a strong referral program, supplemented by a sourcing team that finds people with very specific skills globally.

Oswald attributes Mercor's supply-side success to three factors: a great expert experience (timely pay, transparent compensation, dignified work) that drives referrals, a strong sourcing team that finds niche talent globally, and the compounding effect of both working together to fill spiky demand.

transcript

Oswald Nitski: I probably put it down to three things. The first one is a great expert experience. So experts get paid on time, they get paid well, transparently. Everybody involved in what the expert experience cares deeply about whether or not they're having any challenges and whether or not the work is dignified and well paid and fairly paid. And that is a requirement for a great referrals program because nobody's going to refer their friends or their colleagues to some kind of job that sucks, right? So, everybody caring about expert experience drives a great referral program and additionally a great sourcing team that's able to find people in every corner of the world with very specific skills helps us fill the gaps when we have spiky demand for a specific skill set.

06
Data

The environment-based RL data type is growing fastest at Mercor — these are simulations of apps and rich start states that train agents to use tools — and this shift represents data that looks much closer to what models see in deployment, making it the current frontier of data work.

Oswald identifies RL environments — simulations of apps and computer states for training agents — as the fastest-growing data category. This complex annotation process requires building realistic mocks (like a Salesforce clone) and represents the frontier where data increasingly mirrors deployment conditions.

transcript

Oswald Nitski: The data type that's growing the fastest for us is environments. You might have seen a lot about these RL environments on Twitter. It's kind of a hype term. Every company has a different definition for it. But we are certainly the leader in the category and view it as basically these simulations of apps that you might want your agent to use. And also as per rich start state, which we call the world that is basically representative of all the data you might have on your machine like your laptop. And then we have tasks that train agents how to use those tools to accomplish something that's useful. It's a bit of a complicated annotation process because the agent has to interact with this simulated world. We have to make that start state, which can be hundreds of files, thousands of files. And the shift here is that the data that the models — now the agents — are being evaled and trained on looks a lot closer to what they see in deployment. Right? So, if you want to learn how to use something like Salesforce, you need a pretty high-fidelity mock that acts exactly like Salesforce in your eval and training.

Highlight slides
AI accelerates output, but that makes PM harder✦ from: AI makes the product management job harder, not easier — the challenge shifts from execution to deciding what not to build, as rapid engineering velocity creates pressure to expand surface area chaotically while the real job is fighting to simplify and find the most scalable workflows.Product team response to faster engineering✦ from: AI makes the product management job harder, not easier — the challenge shifts from execution to deciding what not to build, as rapid engineering velocity creates pressure to expand surface area chaotically while the real job is fighting to simplify and find the most scalable workflows.New Bottleneck: Understanding Users✦ from: AI makes the product management job harder, not easier — the challenge shifts from execution to deciding what not to build, as rapid engineering velocity creates pressure to expand surface area chaotically while the real job is fighting to simplify and find the most scalable workflows.RL Environments: Fastest-Growing Data Type at Mercor✦ from: The environment-based RL data type is growing fastest at Mercor — these are simulations of apps and rich start states that train agents to use tools — and this shift represents data that looks much closer to what models see in deployment, making it the current frontier of data work.Why Environments Are the New Frontier✦ from: The environment-based RL data type is growing fastest at Mercor — these are simulations of apps and rich start states that train agents to use tools — and this shift represents data that looks much closer to what models see in deployment, making it the current frontier of data work.Building High-Fidelity Simulations Is Complex✦ from: The environment-based RL data type is growing fastest at Mercor — these are simulations of apps and rich start states that train agents to use tools — and this shift represents data that looks much closer to what models see in deployment, making it the current frontier of data work.
Related episodes