Benchmarks like ARC-AGI, which are deliberately built to be cheap yet currently hard for models, create a regression-to-the-mean effect: once labs target that narrow adversarially-selected weak spot, performance surges upward in a way that looks like a real capability leap but isn't a steady, trustworthy trend.
Responding to the ARC-AGI v1→v2 saturation-then-collapse-then-resaturation pattern, Beth Barnes argues such benchmarks are adversarially selected against current models, which mechanically produces misleading jumps rather than genuine steady progress — part of why Meter avoided that design for time horizon.
transcript
Beth Barnes: with things like RKGI, there is a sort of adversarial selection going on where people are trying to make some benchmark like cheaply subject to the constraint that current models do badly on it. Uh which means you you know you can't use a lot of expensive human labor.