ATRIUMsearch → argument graph
Video · 2026-07-26 · 1h 34m · 6 moments

Why the people building AI can’t tell you what’s next | Dianne Penn (Anthropic)

✦ AI generated

timeline · colored by role

01
Anecdote

Anthropic found its product identity not through a grand strategy, but through small, bottoms-up experiments like Golden Gate Claude that showcased research in an authentic, differentiated way.

Dianne describes how shipping Golden Gate Claude — a quirky 24-hour experiment where Claude obsessed over the Golden Gate Bridge — was an early inflection point that showed Anthropic could build products different from competitors.

transcript

Dianne Penn: I think one of the moments where really we started to get into our groove was shipping things like Golden Gate Claude... when you actually essentially dialed up that feature, Claude would obsess about the Golden Gate Bridge. So meaning in every one of its responses, it would come back and talk about the Golden Gate Bridge... it made us feel like oh we can actually bring new user experiences, showcase our research in a way that's different and authentic to us and in a very startupy like pace. That to me was like one of those like maybe hidden inflection points of we were starting to find our identity that we could build products, build experiences that were different for what our competitors had seen, what was already out there.

extends · 1gives example · 1

02
Claim

Frontier products and frontier models are mutually dependent — Opus 4.5 would not have had its breakthrough moment without Claude Code, and Claude Code would not have had its adoption accelerated without Opus 4.5.

Dianne explains that Opus 4.5's impact came from the combination of a smarter model and a great product experience (Claude Code) — neither alone would have achieved the same result.

transcript

Dianne Penn: I think what was magical about Opus 45 is we also now not just had a model but a vehicle which is like a great product experience like cloud code... you need frontier products in order to have frontier models and for people to feel the magic of frontier models... Opus 45 wouldn't have had that moment without a product like Cloud Code and Cloud Code I think wouldn't have had that type of adoption accelerated without Opus45.

explains mechanism · 1extends · 1supports · 1

03
Mechanism

Being on the exponential curve of AI improvement means adaptability and first-principles thinking matter more than sticking to fixed plans, because emerging capabilities jump discontinuously and unpredictably.

Dianne explains that the scaling law papers show capabilities emerge as discontinuous jumps, not smooth curves, so teams must be adaptable, think from first principles, and be ready to pull forward plans when the model suddenly unlocks something new.

transcript

Dianne Penn: One thing I like to say on the team is most of us weren't like actively working yet when the internet transitioned from this novelty to something that everyone can use and it feels like that's just taking humans... adaptability becomes very important... it's very hard to predict the exact moment or the exact model and so the adaptability of when you're faced with new information how do you then make better decisions versus keeping the same plan... what's actually also interesting in that paper is there are these like very different emerging capability graphs. So for example as you add in more data and you train the models with more compute you essentially see these actually discontinuous emerging capabilities jump... the models go from 1 + one being a thing that it can't calculate to a thing that it can reliably calculate.

explains mechanism · 2extends · 1provides context · 1

04
Claim

Product management in the AI era has shifted: evals are the new PRDs, and PMs must 'sweat the tokens' as much as they sweat the pixels.

Dianne describes how her team's way of driving user value has shifted from writing PRDs to building evals — creating testable, reproducible representations of user pain points that researchers can act on, because the key failures now live in token trajectories, not UI flows.

transcript

Dianne Penn: The way to drive user value is to figure out the right user feedback, the evals, right? That then can be a personification of that user need. So like we do write some product documents and PRDs, but we actually have a saying on the team of evals are the new PRDs... Here, you have to sweat the tokens as much as you sweat the pixels. And so, one activity we have on the team is reading the transcripts and understanding what was the trajectories that failed very deeply to then say was this like a hallucination? was this claw being overconfident.

extends · 1supports · 1

05
Claim

To be an effective manager of AI product teams, you must remain hands-on and spend time actually shipping with the technology yourself, regardless of your seniority.

Dianne argues that managers and even senior leaders need the same hands-on onboarding as early-career PMs — reading user transcripts, shipping features — because you cannot develop good judgment about AI products without experiencing the building yourself.

transcript

Dianne Penn: In order to be good managers of teams and PMs working with this technology, you have to be really hands-on yourself and have spent not just time tinkering but actually shipping with this technology and and and again being in the details and sweating the tokens along with your PMS and your engineer and your teams... even for folks that I hire who have more tenure PM experience, the onboarding plans are exactly the same as somebody who is like more early career and it's around understanding users, reading like consented user feedback, talking to customers... I do feel pretty strongly that if you're a manager, you have to be hands-on. You have to spend a portion of your time actually shipping.

extends · 1

06
Prediction

The most valuable human trait as AI advances is judgment — the accumulation of nuance, experience, and conviction about what to build, which AI systems lack.

Dianne identifies judgment — hard-earned through experience — as the area where human brains will remain most valuable, because there are so many things AI can build and the hard question is which ones to build.

transcript

Dianne Penn: We started to talk about making claude and models better at judgment especially in the last year or so. I think judgment is one and is an area where it's an accumulation of so much nuance and so much experience and these systems haven't experienced as much as humans have and so I think that hard-earned like judgment is a area for product leaders and just generally will continue to be really critical. There are so many things AIs can build. Which one are the things that you know an or like lab should build right a lot of that requires like human judgment persistence so proactivity.

supports · 2

Highlight slides
Related episodes