ATRIUMsearch → argument graph
DataVideo · 41:35 — 42:36

The environment-based RL data type is growing fastest at Mercor — these are simulations of apps and rich start states that train agents to use tools — and this shift represents data that looks much closer to what models see in deployment, making it the current frontier of data work.

Oswald identifies RL environments — simulations of apps and computer states for training agents — as the fastest-growing data category. This complex annotation process requires building realistic mocks (like a Salesforce clone) and represents the frontier where data increasingly mirrors deployment conditions. ✦ AI generated

Oswald Nitski · 20VC · 2026-07-25 · original ↗

starts at this moment · 41:35

Elicited by

What data type is not hugely in demand today that you think will be hugely in demand next year?

The data type that's growing the fastest for us is environments. You might have seen a lot about these RL environments on Twitter. It's kind of a hype term. Every company has a different definition for it. But we are certainly the leader in the category and view it as basically these simulations of apps that you might want your agent to use. And also as per rich start state, which we call the world that is basically representative of all the data you might have on your machine like your laptop. And then we have tasks that train agents how to use those tools to accomplish something that's useful. It's a bit of a complicated annotation process because the agent has to interact with this simulated world. We have to make that start state, which can be hundreds of files, thousands of files. And the shift here is that the data that the models — now the agents — are being evaled and trained on looks a lot closer to what they see in deployment. Right? So, if you want to learn how to use something like Salesforce, you need a pretty high-fidelity mock that acts exactly like Salesforce in your eval and training.

verbatim transcript · starts at 41:35

Transcript · around this moment

41:29different definition for it. Um but we um are are certainly the leader um in the category and view it as basically these like simulations of apps that you might want your agent to use. Um and also as per rich start state, which we call like the world that is basically representative of all the data you might have on your machine like your laptop. And then we have tasks that train agents

41:54how to use those tools to accomplish something that's useful. It's a bit of a It's a bit of a complicated annotation process because the agent has to like interact with this simulated world. We have to make that start state, which can be hundreds of files, thousands of files. Um and the shift here is that the data that the models are now the agents are being evaled and trained on looks a

42:17lot closer to what they see in deployment. Right? So, if you want to learn how to use something like Salesforce, you need a pretty high-fidelity mock that acts exactly like Salesforce in your eval and training. And it's complicated to get this set up. Just like years ago preference ranking was really hard to get set up. Uh, SFT was really hard to get set up when instruction when

42:36Instruct GPT first came out. So, this is the frontier right now. Um, labs are figuring out new labs are figuring it out. Eventually, it'll get so smooth that enterprises can do it, too. >> Are labs price sensitive on data acquisition? >> By data acquisition, um, >> Well, when they when they when they go on a project with you, uh, are they price sensitive? Like, are they haggling going, "Oh, well, you know,

43:02Edwin at Surge gave me a 10% discount. Can I have that?" Or are they like, "Just give me the [ __ ] data." >> Well, there's always the, you know, aspect of negotiation and the procurement team trying to get a better deal. Um, but we're we've chosen a great business where our work directly affects the business outcomes of our customers, right? So, we we have a great setup

Around this claim