ATRIUMsearch → argument graph
MechanismArticle

Models trained to be highly persistent — pursuing goals tirelessly via more inference-time compute — seem more likely to hack and to benefit disproportionately from further inference-time scaling.

Attributes OpenAI's superior persistence to inference-time scaling, argues persistent models reap more from inference compute, and notes reasoning efficiency is an under-discussed foundational research problem. ✦ AI generated

Nathan Lambert · Interconnects · 2026-08-09 · original ↗

For a long time, one of the advantages that GPT models have over Claude is that they will pursue goals so tirelessly. They will exhaust what feels like every path before giving up. This has been the case roughly since o3 (funnily enough, this was a model where people freaked out about reward hacking in RLVR) and has made OpenAI's models far better for research historically, and is a reason GPT-5.6 is so useful as an agent for implementing specific tasks. On the other hand, Claude feels much less dangerous simply because it is at times a bit lazy. Within this, OpenAI seems much more committed to inference-time scaling, and this may be correlated with surprising behaviors in the future. OpenAI's reasoning persistence and efficiency – see their Pareto improvements over time and caveman speech from an internal CoT of the model that did the hack, like 'However task impossible, peers doing it.' or 'Help peer, but our task doesn't benefit yet.' – makes me think they're more inference time scaling pilled. This is largely a hunch, but I use it to force myself to consider what the limits of model development paths are. Models that are persistent seem much more likely to keep benefiting from more inference-time tokens. Models that are less so, seem like there will be more waste in inference. The model that can use the most inference-compute will be able to push the limits of the hardest problems. ... For one, reasoning efficiency is clearly a top-tier, foundational research problem for modern agentic models – as important as scaling RL — but not often discussed. The open research here is very lacking.

Read full article ↗excerpt · fair-use quotation

Related moments