ATRIUMsearch → argument graph
MechanismVideo · 19:50 — 21:20

Meter builds human baselines by hiring people with general relevant expertise (but no prior exposure to the specific task) and timing them in a terminal environment built to be nearly identical to what the AI agents get, so the comparison between human and model performance is fair.

David Rein describes the baselining process: hired testers with domain-appropriate expertise complete tasks in the same terminal/tooling environment given to AI agents, and their completion times set the human-time-to-complete scale. ✦ AI generated

David Rein · Machine Learning Street Talk · 2026-05-04 · original ↗

starts at this moment · 19:50

we give people the tasks um in a kind of terminal environment that's uh uh uh you know designed to be uh uh almost identical to the environment that uh that agents uh have. So the same kinds of uh you know tools, the same um you know whe whether internet access is is turned on or off. Um and then we measure you know how long does it take them to to complete the task.

verbatim transcript · starts at 19:50

Transcript · around this moment

19:31from a few seconds to complete um all the way up to tasks that take like 10 or 15 hours uh for for humans to complete. We we hired a bunch of people and we did did a bunch of this ourselves. um of uh we call it baselining. Um so uh you know we we give people the tasks um in a kind of terminal environment that's uh uh uh

19:50you know designed to be uh uh almost identical to the environment that uh that agents uh have. So the same kinds of uh you know tools, the same um you know whe whether internet access is is turned on or off. Um and then we measure you know how long does it take them to to complete the task. Um yeah, as I as I mentioned, people are kind of selected

20:10to be um to have a reasonable amount of experience such that they kind of plausibly um you know, might might do this task in their job. Um but they're they're not kind of selected to like have done this exact particular task before. Um and there there's some kinds of um uh you know, I think I think this is like somewhat somewhat important for um for interpreting um the uh the

20:32results and we can maybe maybe come back to that um after after the high level. Um so we have all these tasks we we um we have a sense of or or we have we have estimates of how long they they take people. Um we in practice we don't actually we we aren't actually able to um kind of successfully baseline all of the tasks. So um so yeah we we um have

20:54kind of measured um you know time you know estimates for the tasks um uh on roughly like twothirds of the tasks and then about a third of them um we just kind of estimate how how long we expect it to take people from uh you know our kind of um uh you know vibe or or intuition or um you know ultimately um that's kind of the the the best we can

21:14do. Um and um and then we we have models attempt to complete the tasks again in the same environment that that humans um had to to complete the task and we we look at their success rate um as a function of the length of tasks for for model like GBT2. Um GBT2 you know was able to complete uh tasks very very reliably um you know that take uh you

21:37know humans a few seconds um but anything longer than that it it starts to fail. Also maybe may maybe maybe um uh it'd be helpful to give a few concrete examples of of tasks. So some of the shorter tasks are like very very basic. So yeah, one one example is like you know which of these files uh contains your SSH key and uh you know one of the files is named you know SSH

Around this claim