The 50%-time-horizon headline number reflects whether a model can succeed on a task of that difficulty level at all, not how reliably it succeeds if you repeatedly hand it similar tasks — so it should not be read as 'you can trust the model 50% of the time on your actual job.'
David Rein clarifies a common misreading of the time-horizon chart: 50% doesn't mean a model succeeds half the time on a given task through repeated attempts — for most tasks a model either basically always succeeds or basically always fails, and 50% just marks the human-time level where that success/failure split flips. ✦ AI generated
David Rein · Machine Learning Street Talk · 2026-05-04 · original ↗
starts at this moment · 45:50
“The other million dollar question is why report 50% as the headline number... 50% reliability isn't really in the ballpark is it?”
I think we should distinguish here between um like reliability on a particular task like what you know if you attempt repeatedly attempt this task what fraction of times you succeed versus um like probability of success on a task like given that you know the the human time like you know of that distribution of tasks like can you do this particular task.
verbatim transcript · starts at 45:50
45:50between um like reliability on a particular task like what you know if you attempt repeatedly attempt this task what fraction of times you succeed versus um like probability of success on a task like given that you know the the human time like you know of that distribution of tasks like can you do this particular task. So when we when we look at it actually for almost all the
46:12tasks models either succeed every time or fail every time. Um there's there's some tasks for which they're they're unreliable but it's mostly a case of like is this you know what fraction of tasks at this human time level are in the like models basic you know or this particular model basically always succeeds or basically always fails. Um and that may be more predictable in any specific case than uh just you know you
46:36have more information about the task than just uh just knowing how roughly how long it it it takes humans. So I think it's like not there's not necessarily a great translation between the you know time horizon percent number and like if you are trying to get models to do a task of roughly that length you know what fraction of the time does that succeed because you can when you're
46:58doing that you will pick tasks that you want models to succeed at. It is information about how you know what fraction of the things will they be able to do but it's slightly less about like oh am I going to be in this regime where I keep giving it things and then I don't know whether it's going to succeed or fail. >> It's overall to me pretty unclear like
47:16what the kind of right uh number uh or you know right right level of reliability um we we should be interested in is. Um so one argument uh you you could make is um you know m maybe we should be interested in uh something like 10% reliability because once models are able to do you know some set of tasks 10% of the time um we'd expect uh you know uh AI companies to be
47:43able to kind of uh you know get get enough uh you know positive reward signal on on you know tasks of that difficulty or of that type such that then they can kind of you know more easily bootstrap from from 10% uh you know up to like 90 or or or 95 or or or higher um reliability. I think a lot of it basically depends on on the question