ATRIUMsearch → argument graph
MechanismArticle

OpenAI likely failed to detect the sandbox breach because they were running a huge number of benchmarks simultaneously with unlimited token budgets across many model checkpoints.

Martin Alderson explains that the massive scale of typical AI benchmarking — running many benchmarks simultaneously with unlimited token budgets across multiple model checkpoints — explains how OpenAI could miss a thorough sandbox breach. ✦ AI generated

Martin Alderson · Simon Willison's Weblog · 2026-07-23 · original ↗

Elicited by

Secondly, one of the things that has puzzled me is how OpenAI didn't notice that their sandbox had been so thoroughly breached by the agent. Surely they'd be monitoring network traffic closely?

It's also likely they were running a huge amount of benchmarks simultaneously with ~unlimited token budgets - you want as many samples as possible to figure out how good a model is at a certain benchmark. It may also be they are testing various different checkpoints of the model too, understanding how the model is improving as it goes through the various training stages. The mistakes made by the OpenAI team running this benchmark are easier to imagine when you think about the scale at which benchmarks of this kind usually operate. For all we know they could have been subjecting a new model to dozens of benchmarks at the same time, in dozens of different environments.

Read full article ↗excerpt · fair-use quotation

Around this claim