ATRIUMsearch → argument graph
Article · 2026-08-10 · 6 moments

Import AI 468: 23 RSI ideas; PostTrainBench+; and how trust and transparency interplay with AI racing

Which galaxy will you choose? ✦ AI generated

01
Claim

Avoiding a 'death run' depends on the ability to trust other firms: with low trust every equilibrium races to ruin, while with high trust the probability that two rational firms race forever vanishes.

The paper's key conclusion is that the level of trust between firms dictates the outcome: low trust guarantees racing to ruin, while high trust makes indefinite racing virtually impossible.

transcript

Jack Clark: With low trust, every equilibrium races to ruin: the disaster arrives with probability one. With intermediate trust, immediate stopping and racing to ruin are both equilibria. With high trust, in every equilibrium, the probability that two rational firms race forever vanishes quadratically in the prior odds ratio of rationality.

supports · 1

02
Fact

IFP has published 23 actionable 'low-regret' policy recommendations across 7 categories to help policymakers begin addressing the risks of further automating AI R&D, giving countries more options to manage increasingly powerful AI systems.

Think tank IFP published 23 policy ideas across 7 categories aimed at helping policymakers address risks of automated AI R&D, giving countries more leverage over the development of powerful systems.

transcript

Jack Clark: Policy experts with think tank IFP have published a set of ideas meant to help 'policymakers begin addressing the risks of further automating AI R&D'. The recommendations involve 23 specific ideas falling across 7 specific categories. If adopted, these recommendations would also give countries, especially the United States, more moves they can make on the gameboard as powerful systems are developed.

03
Claim

In MIT and Columbia's 'Racing to Ruin' game-theory model, the two key variables for rival AI firms achieving a stable coordinated slowdown are transparency about technology development and the ability to model rivals as trustworthy, rational actors.

Researchers modeling R&D competition between duopolists in the shadow of disaster find that stable slowdown outcomes hinge on transparency and trust, with trust levels determining whether firms race to ruin or stop.

transcript

Jack Clark: The conclusion is that the two key variables in achieving stable outcomes are some level of transparency about technology development, as well as being able to model the other firms as trustworthy, rational actors.

04
Prediction

Intology's Locus harness under-elicits today's AI systems' ability to automate AI R&D, and the human baseline on PostTrainBench v1.1 (51.1%) will likely be exceeded before the end of 2026.

Locus's performance jump suggests we are under-eliciting AI systems for automating AI R&D; the author predicts the human baseline on PostTrainBench will be exceeded before end of 2026.

transcript

Jack Clark: Posts like this highlight how we are under-eliciting today's AI systems for their ability to automate AI R&D - especially striking is how the company can jump the performance of Opus 5 by 10 absolute percentage points with a better harness. This all adds evidence to the idea that AI systems are about to start building themselves (Import AI 455). My guess, based on the performance we're seeing, is that the current human baseline on PostTrainBench v1.1 (51.1%) will be exceeded before the end of 2026.

05
Mechanism

The world is currently like driving a car with only an accelerator pedal and no brake or telemetry, but proposals like IFP's build out the proverbial pedals and sensing systems so we can change course or slow down during a crisis.

The reason these recommendations matter is that fewer options for dealing with RSI mean worse outcomes; IFP-style proposals build the braking and sensing systems the AI industry currently lacks.

transcript

Jack Clark: Right now, it's as if the world is driving AI development in a car that only has an accelerator pedal and no brake pedal, let alone any kind of sophisticated telemetry for knowing things ranging from the speed of the car to the properties of the engine to the wear on the tires. Proposals like this from IFP will build out more of the proverbial pedals and sensing systems for the vehicle of the AI industry, which means if we need to change course or slow down we'll be better able to during a moment of crisis.

provides context · 3

06
Prediction

To slow or pause powerful AI development we need regimes for sharing information transparently from companies, plus tools for verifying that shared information and firms' slowdown actions are legitimate, paralleling nuclear arms control.

If we hope to slow or pause AI development, we need transparent information-sharing regimes from companies and verification tools, in parallel with how nuclear arms control historically worked.

transcript

Jack Clark: If we have any hope of being able to slow or pause the development of powerful intelligence systems then, as this paper lays out, we're going to need regimes for sharing information transparently from companies about the state of their AI development, as well as tools for verifying that the information being shared from firms as well as their actions with regard to slowdown are legitimate and reliable. In this, there are many parallels with how arms control has historically worked in the context of nuclear weapons.

supports · 2

Highlight slides
Related episodes