ATRIUMsearch → argument graph
Article · 2026-08-17 · 6 moments

Import AI 469: Science AI; RSI simulator; and Zuck's technological pessimism

The new frontier of AI is developing capable autonomous researchers ✦ AI generated

01
Data

Current frontier models cannot beat DiG-bench, with only Opus 5 and Fable 5 reaching Tier 7, suggesting AI still struggles with creative discovery compared to humans.

Despite being difficult, games in DiG-bench are beatable by humans, but today's best AI models struggle significantly, indicating a gap in discovery capabilities.

transcript

author: Overall, this seems really hard! ... some frontier models are already capable of some fairly impressive feats of discovery, but still struggle compared to humans (for instance, a 20% success rate on Tier 7 is pretty poor compared to the fact individual humans were able to get 100% on the tests).

02
Claim

DiG-bench measures AI's ability to discover hidden rules in novel environments through exploration, which serves as a proxy for creative intuition.

DiG-bench is a benchmark of 70 games designed to test whether AI can uncover hidden rules through interaction, measuring a core capability for creativity.

transcript

author: How well can AI systems figure out the rules of their environment through exploration and curiosity, versus being fed them? That's an important question for better understanding the intuitive and creative capabilities of AI systems and it's one being asked by DiG-bench (Discovery in Games), a new benchmark of 70 games "designed to map the surface of discovery in well-controlled interactive systems".

03
Example

Games that simulate recursive self-improvement can help build better intuitions about AI development and its existential implications.

Paradigm Research's browser game allows people to experience the challenges of building AI systems capable of recursive self-improvement, helping build intuition about this important technology.

transcript

author: Here's a fun game from the folks at Paradigm Research which aims to simulate what it's like to run a company building AI systems which become capable of recursive self-improvement... Developing better intuitions about recursive self-improvement is of existential importance to us all; games like this help make it easier for us to reason about this technology and the labs building it.

05
Claim

Inherent's Faraday model can learn to replicate research experiments, with capabilities that compound and could enable AI to advance its own state of the art.

Faraday, a small supervisory model trained on top of frontier models, demonstrates research capabilities that could eventually allow AI systems to design their own experiments and improve themselves.

transcript

Inherent: The skills that allow Faraday to fill in vaguely-specified details may be the very same skills that would allow it to advance the state of the art by designing its own experiment... One might hope that a single post-trained outer agent can track the frontier as better models are released, at least over some time period.

06
Claim

Zuckerberg's essay ignores the fundamental question of whether systems capable of superhuman invention will actually serve individual empowerment rather than fundamentally altering the balance of power.

The author argues that Zuckerberg's vision of proliferating AI to empower individuals fails to address whether truly capable invention systems would actually work on behalf of less capable humans.

transcript

author: The missing question: The part of this essay I understand the least is Zuckerberg's co-mingling of AI systems capable of invention with individual empowerment... Surely this is the key question? I am not suggesting that superhuman invention guarantees some kind of malign entity that is independent from people. Rather I am suggesting that it's hard to reconcile a system capable of superhuman invention with something that doesn't fundamentally alter the balance of power in the world in ways that are confusing and hard to reason about.

Highlight slides
Related episodes