ATRIUMsearch → argument graph
Video · 2026-07-26 · 1h 34m · 6 moments

Why the best product leaders are building for 2028 | Dianne Penn (Anthropic)

✦ AI generated

timeline · colored by role

01
Anecdote

Rapid, bottoms-up, cross-team experimentation — not grand strategy — was the mechanism through which Anthropic discovered its product identity.

Dianne recounts the Golden Gate Claude experiment — a quirky interpretability demo shipped in 24 hours by volunteers — as a hidden inflection point where Anthropic learned it could differentiate through fast, authentic product experiences.

transcript

Dianne Penn: The entire experience actually we spun up on our claude.ai website within 24 hours and that took like engineering, product, design, our research teams all working together and we were really really proud of it. I think it maybe reached only 2,000 people to be honest. But it made us feel like oh we can actually bring new user experiences, showcase our research in a way that's different and authentic to us and in a very startupy pace. That to me was like one of those maybe hidden inflection points of we were starting to find our identity that we could build products, build experiences that were different for what our competitors had seen, what was already out there.

gives example · 1

02
Mechanism

Identifying long-form code generation as a training focus for Opus 3 gave Anthropic its first competitive differentiation against OpenAI.

Dianne explains that in 2023 no one associated Anthropic with coding until she noticed users writing long-form code with models, leading to a small training change for Opus 3 that became a major differentiator.

transcript

Dianne Penn: In 2023 when I started nobody said anthropic and claude and coding in the same sentence. I think competitor models like GPT4 at the time was used a bit for coding but it was one of many use cases. And one thing that for example I saw was people are starting to use these models not just for code autocomplete but actually writing long form code and is that an opportunity for us to train Opus 3 to be better at and it ended up being a relatively smaller change from a training perspective but it ended up helping us differentiate in the early days competitively for users and actually bring a lot of the very early Claude enthusiasts and developers because we were providing a value that they didn't really think was possible at the time.

explains mechanism · 1provides context · 1

03
Claim

Breakthrough AI moments require both a frontier model and a frontier product — each accelerates the other and neither reaches its full potential alone.

Dianne describes the symbiotic relationship between Opus 4.5 and Claude Code: the model needed the product to deliver its magic, and the product needed the model's intelligence to drive adoption.

transcript

Dianne Penn: Opus 45 was definitely another large moment. I think what was magical about Opus 45 is we also now not just had a model but a vehicle which is like a great product experience like Claude Code. One thing we say a lot on the team is you need frontier products in order to have frontier models and for people to feel the magic of frontier models. I think Opus 45 wouldn't have had that moment without a product like Claude Code and Claude Code wouldn't have had that type of adoption accelerated without Opus 45.

gives example · 3

04
Data

As AI models scale, they exhibit discontinuous emerging capabilities — abilities like calculation that suddenly appear — making it inherently unpredictable what new skills a model will gain.

Dianne references the original scaling law papers to explain that model capabilities do not improve smoothly — they jump discontinuously — requiring evaluation systems to detect what the model can suddenly do.

transcript

Dianne Penn: There's some really interesting graphs in the original scaling law papers. I think folks are very familiar with the scaling loss in the lens of as you add in more compute and data what's called loss aka the loss from next token prediction goes down. And so it's a very smooth linear curve of like the models get more intelligent as you scale them up. What's actually also interesting in that paper is there are these very different emerging capability graphs. And so for example as you add in more data and you train the models with more compute you essentially see these actually discontinuous emerging capabilities jump. So the models go from 1+1 being a thing that it can't calculate to a thing that it can reliably calculate. And so these emerging capabilities, like some nature of predictability is not necessarily everyone knows the exact moment — you need the evals to be able to assess that — has actually always been a part of how this technology works and also what makes things like safety harder because unless you have the eval, unless you have the systems to test, these jumps might actually happen and you don't know.

05
Claim

In AI product management, evals have replaced PRDs as the primary tool for defining and measuring user value.

Dianne explains that for research product managers at Anthropic, translating user feedback into standardized model evaluations has become more central than writing product requirement documents.

transcript

Dianne Penn: I think one example is I think of a product manager as I own product strategy and delivering user value, but I demonstrate day-to-day by writing a PRD or writing a product vision doc. And for my team as research product managers, the way to drive user value is to figure out the right user feedback, the evals, right? That then can be a personification of that user need. So like we do write some product documents and PRDs, but we actually have a saying on the team of evals are the new PRDs, right? Because in order to deliver that user value it's not that exact artifact that people used to write in the last one to two decades, it's a new way of working.

extends · 1

06
Prediction

Human judgment — hard-earned through lived experience and nuance — remains the most valuable asset as AI capabilities advance.

Dianne identifies judgment, persistence, and proactivity — traits built through accumulated human experience — as the areas where people remain indispensable even as models grow more capable.

transcript

Dianne Penn: We started to talk about making Claude and models better at judgment especially in the last year or so. I think judgment is one and is an area where it's an accumulation of so much nuance and so much experience and these systems haven't experienced as much as humans have and so I think that hard-earned judgment is an area for product leaders and just generally will continue to be really critical. There are so many things AIs can build. Which one are the things that an AI lab should build right? A lot of that requires human judgment, persistence, proactivity. These are all traits that are beyond just general capabilities but just behaviors and characteristics of people at that level of how do you get to the best solutions, how do you create the best experiences. So I think those types of traits are actually the tactile traits that I think will continue to be important.

Highlight slides
Related episodes