The capability of a model is now a function of how much test-time compute you invest in it, not a fixed property, and existing evaluation frameworks (preparedness frameworks, responsible scaling policies) do not account for this.
Noam argues that traditional safety and capability evaluations treat model performance as a static number, but with modern models capable of sustained reasoning over weeks, capability scales with compute budget — a question the current frameworks ignore.
transcript
Noam Brown: The preparedness frameworks and responsible scaling policies, they don't really account for the amount of test time compute. They just say, okay, well, what's the capability of the model? The problem is we're in a world now where the capability of the model is a function of how much money you put into it, basically. If you give it a budget of $10,000, it can do a lot more than what it can do with a budget of $10. Give it a budget of $10 million, it could do even more. And so at what budget should you evaluate these models? The policies that exist today don't really address that question.
extends · 1