ATRIUMsearch → argument graph
PredictionVideo · 101:09 — 102:39

Full autonomous AI self-improvement this year is unlikely but not low-probability enough to rule out, and the plausible path runs through accelerating time-horizon gains on easily-verifiable tasks generalizing further, plus compounding low-hanging-fruit improvements in RL training environments, compute efficiency, and scaffolding that together speed up AI R&D itself.

Asked to unpack her claim on another podcast that recursive AI self-improvement could arrive within two years, Beth Barnes assigns it a low but non-negligible single-digit percent chance this year, sketching a pathway through accelerating time-horizon trends, better post-training environments, and compute-efficiency gains compounding into faster automated R&D. ✦ AI generated

Beth Barnes · Machine Learning Street Talk · 2026-05-04 · original ↗

starts at this moment · 101:09

Elicited by

You said that AI could autonomously self-improve within as little as 2 years and maybe even shorter timelines were hard to rule out. Could you walk through the concrete sequence of steps that could lead to that kind of recursive self-improvement?

I'm like this seems very unlikely to happen this year but it's not you know not unlikely enough to rule out and I think that basically looks like maybe we would see accelerating trend in time horizon on like like easily hill climbable tasks. And it turns out that was actually, you know, a much more general capability.

verbatim transcript · starts at 101:09

Transcript · around this moment

101:09this seems very unlikely to happen this year but it's not you know not unlikely enough to rule out and I think that basically looks like maybe we would see accelerating trend in time horizon on like like easily hill climbable tasks. And it turns out that was actually, you know, a much more general capability. And there was just, you know, a bit of something you needed to do to sort of

101:32like elicit uh it on on these less less hel climbable tasks. But sort of, you know, fundamentally they they are using the same capabilities in a model. It was just sort of, you know, what what you trained on that was affecting the difference we're seeing. Um then this [clears throat] is leading to Yeah. like you automate and accelerate a bunch of AR and DD. So I think I think there's

101:56you know there are a lot of lowhanging fruit even of things that we already know that you could do this and it would improve model performance. Um and it's just you know it doesn't require new breakthroughs and it's just kind of labor intensive to do. So just making much much better post-training environments and really you know crafting them to to teach all the new abilities that you want and and I think

102:17you can probably improve compute efficiency a bunch with you know again just like applying a bunch more labor to like making all your kernels more efficient and also you know doing the right kind of routting between different models or or or other things like that. There's like lots of ways in which uh how we're using compute is not optimized. So you could potentially get a bunch of you know sort of like the

102:37equivalent of much more compute scaling out of that. And then you know the hypothesis is like also like scaffolding and training the models to use particular scaffolding and sort of like use memory and retrieval in the right way like it seems kind of obvious that like you know if you if you sort of really had all the right training data and you have a you know you have a

102:58transformer and and it can kind of fill its context with with different things and and take stuff in and out. It can do a pretty, you know, good job of something looks like the sort of continual learning or or building up understanding if you've got massive massive context window and enough, you know, you've actually got quite a lot of bits in there to sort of be adding things about what you've been

Around this claim