ATRIUMsearch → argument graph
ClaimVideo · 93:23 — 94:53

Forecasting is the only completely renewable source of hard questions with guaranteed correct answers, which makes it the best available evaluation for AI capability now that human experts can no longer write evals the models can't already ace.

Dan Schwarz argues forecasting uniquely provides limitless hard questions with future ground-truth, making it the 'ultimate eval' for frontier labs as human experts lose the ability to author unbeatable benchmarks. ✦ AI generated

Dan Schwarz · The Cognitive Revolution · 2026-07-07 · original ↗

starts at this moment · 93:23

Elicited by

I know you said that you can't share how they are thinking about it, but I wonder if you could give us some thoughts on how you think they should be thinking about it.

I think uh human experts, doctors, lawyers, engineers, financers, whoever who are trying to make evals to try to produce data for the frontier labs are finding that they are not smarter than the things being trained anymore. And so if you can produce something that has a correct answer, the model's already going to figure out that correct answer. You need something where there's a correct answer and the model can't figure it out. Forecasting, I think, is the only completely and utterly renewable source of this.

verbatim transcript · starts at 93:23

Transcript · around this moment

93:03three weeks in the future. And so you basically have a completely limitless set of extremely hard basically impossible questions where you get exact ground truth. And there is no other email like this. There is if you want to say improve a coding harness, you just need to have more and more hard coding problems that are not in the training data for which you can say this is

93:23definitely the correct answer so that you can do some training on it. And that's hard. I think uh human experts, doctors, lawyers, engineers, financers, whoever who are trying to make evals to try to produce data for the frontier labs are finding that they are not smarter than the things being trained anymore. And so if you can produce something that has a correct answer, the model's already going to figure out that

93:43correct answer. You need something where there's a correct answer and the model can't figure it out. Forecasting, I think, is the only completely and utterly renewable source of this. And again, this kind of is connected to forecasting as the, you know, the the kind of Elon Musk tweet the quips that forecasting is like the ultimate measure of intelligence. If you zoom out and think about it in one perspective, it

94:02is. Again, I think coding intelligence, AI, R&D intelligence, interpersonal intelligence are pretty darn important. I wouldn't say that forecasting is truly ultimate intelligence, but it is to some degree the ultimate eval. And I think this is something that Frontier Labs like Canon should be paying attention to. Do you have a sense for how much irreducible chaos there is? Like in the limit of super intelligence, how much visibility

94:31do we get into the future? >> Yeah, that is a great question. Uh David Mannheim actually just uh posted a tweet that has a graph that has a claim to estimate this that I was meaning to dig into because I don't know where those lines are coming from. Um I would say without evidence behind this my sense is very Ukowskian like I think there is a lot of detail in reality that is far

94:54beyond the human mind to understand and as you approach more sophisticated intelligence you will start seeing a lot of patterns um and then the point of trying to produce you know voxal perfect weather three weeks in the future is further away than people think. I think there's quite a lot of room. Um, human super forecasters don't tend to agree with me on this. They basically think that what they're doing is somewhat near

95:18optimal and any sort of accuracy improvements you're going to get over them is going to be tiny and like hard to understand. And I think that's just because we only really understand human intelligence. And when you kind of just zoom out from an information theory perspective, from like a Kolamor of complexity, like just modeling the world as bite strings, uh, the AI overlords will eventually start to figure out

Around this claim