ATRIUMsearch → argument graph
MechanismVideo · 14:37 — 16:07

Benchmarks like ARC-AGI, which are deliberately built to be cheap yet currently hard for models, create a regression-to-the-mean effect: once labs target that narrow adversarially-selected weak spot, performance surges upward in a way that looks like a real capability leap but isn't a steady, trustworthy trend.

Responding to the ARC-AGI v1→v2 saturation-then-collapse-then-resaturation pattern, Beth Barnes argues such benchmarks are adversarially selected against current models, which mechanically produces misleading jumps rather than genuine steady progress — part of why Meter avoided that design for time horizon. ✦ AI generated

Beth Barnes · Machine Learning Street Talk · 2026-05-04 · original ↗

starts at this moment · 14:37

Elicited by

I mean what do you think about that?

with things like RKGI, there is a sort of adversarial selection going on where people are trying to make some benchmark like cheaply subject to the constraint that current models do badly on it. Uh which means you you know you can't use a lot of expensive human labor.

verbatim transcript · starts at 14:37

Transcript · around this moment

14:37some benchmark like cheaply subject to the constraint that current models do badly on it. Uh which means you you know you can't use a lot of expensive human labor. So it has to be something that's either automatically checkable or that you can like create with kind of cheap human labor. And then but once you've selected on like those things and on models being bad at it, this is now you

14:57know there's like regression to the mean type thing where it is much more likely that future progress then gives you a like big surge upwards on on that both because like you know now labs will create a bunch of bunch of synthetic data targeting your benchmark but also just because you selected this weird example where it's like easy for humans or it's automatically generatable or checkable but somehow like models aren't

15:21good at it yet or you know labs haven't started training on it yet. So I think that's part of what we were trying to do with time horizon was not do that like not adversarially select against what models can currently do because we think that will not give you a nice trend. Uh whereas if you can sort of define a distribution of tasks in some more first

15:38principles way you would be more likely to get a steady um progress because you're not getting this sort of regression to the mean effect. >> Yeah. Frantois has this idea that there is a kind of there's a gap between the kind of intelligence for one of a better word that that AIs have and that humans have and we can adversarily select a bunch of tasks to kind of you know to to

15:58highlight that gap but we should talk about the timeline stuff. I think we'll come back to intelligence later. So um Dan cockatel he said that the timelines report that that you folks have created is probably the single most important piece of evidence about timelines right now. So it it should be um front and center in in policy discussions and and so on. And for listeners who have only

16:20kind of seen the chart but you know they've not really read the paper, they don't understand it. Can you just go through it from a high level? I mean it's been revised over time uh you know how did you do the task selection? You know how did you do the the human baselines? How do you do the agent harness? Like all of that kind of stuff. I guess the yeah the the place to start

Related moments