ATRIUMsearch → argument graph
Video · 2026-07-26 · 1h 34m · 6 moments

Why AI is going vertical (again) | Dianne Penn (Anthropic)

✦ AI generated

timeline · colored by role

01
Anecdote

Anthropic's bottoms-up culture, where engineers and designers donated their time to build quirky experiments like Golden Gate Claude in 24 hours, is the foundation of the company's identity and ability to differentiate from competitors.

Dianne recounts how Golden Gate Claude — a playful feature built in 24 hours by engineers, designers, and researchers donating their time — exemplified Anthropic's bottoms-up, mission-driven culture and became a hidden inflection point where the company found its differentiated identity.

transcript

Dianne Penn: And so it was like really quirky and we we we very much wanted to in that situation just bring that user bring bring it to the masses and bring it to people who are starting to use claude and uh so the entire uh experience actually we spun up on our cloud.ai I website within 24 hours and that took like engineering, product, design, uh our like research teams all working together and we were really really proud of it. I think it maybe reach only 2,000 people to be honest. Uh but it it made us feel like oh we can actually bring new user experiences, showcase our research in a way that's different and authentic to us and in a very startupy like pace. That to me was like one of those like maybe hidden inflection points of we were starting to find our identity that we could build products, build experiences that were different for what our competitors had seen, what was already out there. And I think that obviously labs, clog code, etc. Like we then started to identify ourselves as what we actually think the world uh how to think about AI, how to bring that closer to the public. Um but it was a very bottoms up culture. And so that entire experience was very bottoms up. I see engineers, I see uh designers donating time to work on. Um, and so I I like to always use that as example of like what the day early days were like, but the culture and and and the values have very much I think stayed the same since those early days.

explains mechanism · 1

02
Mechanism

Frontier models and frontier products create a mutually reinforcing flywheel — Opus 45 would not have had its moment without Claude Code, and Claude Code would not have seen accelerated adoption without Opus 45.

Dianne traces two inflection points: Opus 3 proved Anthropic could build a frontier model and identified coding as a differentiator; Opus 45 then combined with Claude Code in a virtuous cycle where the model's intelligence and the product's experience amplified each other.

transcript

Dianne Penn: Um Opus 45 was definitely another large moment. I think what was magic about magical about Opus 45 is we also now not just had a model but a vehicle which is like a great product experience like cloud code. Um one thing we say a lot on the team is you need frontier products in order to have frontier models and for people to feel the magic of frontier models. And I think you know we felt the magic of cloud code for very for for uh for many months before that. Uh but the fact that the model essentially got to a level of intelligence where at a very broad level users can experience both frontier intelligence in new use cases allow it to run things end to end in an agent manner. I think that was the inflection. It was actually both. I I think Opus 45 wouldn't have had that moment without a product like Cloud Code and Cloud Code I think wouldn't have had that type of adoption accelerated without Opus45.

03
Claim

AI capabilities emerge in discontinuous jumps rather than smooth improvements — models can go from unable to do something to reliably doing it, which makes safety harder and demands that organizations have eval systems in place to detect these jumps when they happen.

Dianne explains that while scaling loss follows a smooth curve, emerging capabilities jump discontinuously — models unpredictably gain abilities, which makes safety testing harder and requires robust eval systems to detect what the model can now do.

transcript

Dianne Penn: I think um there's some really interesting graphs in the original scaling law papers and I think folks are very familiar with the scaling loss in in the lens of um as you add in more compute and data what's called loss aka the loss from next token prediction uh goes down. And so it's a very smooth linear curve of like the models get more intelligent as you scale them up. What's actually also interesting uh in that paper is there are these like very uh different emerging capability graphs. And so for example uh as you add in more data and you train the models with more compute you essentially see these actually discontinuous emerging capabilities jump. So the models go from 1 + one being a thing that it can't calculate to a thing that it can reliably calculate. And so these emerging capabilities, this like some nature of like predictability is is is not necessarily everyone knows the exact moment like you need the evals to be able to assess that has actually always been a part of uh how this technology works and also what makes like things like safety harder because unless you have the eval unless you have the systems to test um these jumps might actually happen and you don't know.

04
Mechanism

For research product managers at Anthropic, evals are the new PRDs — the primary artifact for defining user value has shifted from documents describing what to build to reproducible tests that measure whether the model has improved on specific user pain points.

Dianne describes how the PM role has fundamentally changed: instead of writing PRDs, AI product managers translate user feedback into evals — structured tests that capture the exact failure trajectory (hallucination, overconfidence, tool-calling error, etc.) so researchers can measure improvement across model generations.

transcript

Dianne Penn: I think you think of a product manager as I own product strategy and delivering user value as but I demonstrate day-to-day by writing a PRD or writing a product vision doc. And for for my team as like research product managers, the way to drive user value is to figure out the right user feedback, the evals, right? That then can be a personification of that user need. So like we do write some product documents and PRDs, but we actually have a saying on the team of evals are the new PRDs, right? because in order to deliver that user value uh it's not that exact artifact that people used to write in the last like one to two decades it's a new way of working and so the first think principal thinking would be let me figure out what is the thing I should do to achieve my goals rather than here is a set of activities that I've done and therefore I will continue to do.

explains mechanism · 1

05
Claim

AI product managers and leaders at every level must be hands-on builders who personally ship with the technology and sweat tokens alongside their teams, or they cannot develop the judgment to make good decisions.

Dianne argues that tenure offers no shortcut: senior PMs and managers need the same hands-on onboarding as new hires because you cannot evaluate what good looks like in AI products without building yourself, and she personally carves out time to own one to two work streams per model cycle to stay grounded.

transcript

Dianne Penn: One thing that I think I feel pretty strongly about is in order to be good managers of teams and PMs working with this technology, you have to be really hands-on yourself and have spent not just time tinkering but actually shipping with this technology and and and again being in the details and sweating the tokens along with your PMS and your engineer and your teams. And so even for folks that I hire who have more tenure PM experience, the onboarding plans are exactly the same as somebody who is like more uh early career and it's around understanding users, reading like consented user feedback, talking to customers. I think there's something around uh being able to like understand what to do with this, what what good looks like and having developed that in a very hands-on manner. That's important. Um it's not necessarily easy for someone to uh agree or be able to see what a what a good or great AI product or AI feature could look like if they haven't kind of experienced building themselves. Um, so I think I think there is a I I I do feel pretty strongly that like, you know, if you're a manager, you have to be hands-on. You have to spend a portion of your time actually shipping. You you have to kind of walk in the shoes of your teams. uh and and that's I I always try to carve out a portion of time uh to to actually like own one to two work streams when we have models in order to keep like keep my theory of mind, keep my sense of how the models are moving, how quickly it's improving uh uh so I can help the team make make decisions and and make better decisions.

extends · 1

06
Claim

AI does not eliminate the need for product managers — it increases it, because the hard part is no longer building but deeply understanding users and deciding what to build, and that requires people who are relentlessly user-centric, hands-on, and able to translate user needs into actionable form.

In her closing statement, Dianne directly addresses the question of whether PMs are still needed in the AI era and argues the opposite: as building becomes easier, the binding constraint becomes user empathy, deep curiosity, and the relentless work of understanding what people are trying to accomplish — which is the core of product management.

transcript

Dianne Penn: I think the role of people who are user centric who go into the details of understanding what users are trying to accomplish bubbling that up in an actionable manner and doing the relentless work to do that like that to me is a core of a product person and I actually think we need more of that. I think we are becoming very technology layered driven and actually to make that impactful it's you have to go deep you have to be curious you have to be super hands-on and those are things that I think are also traits that have I think helped anthropic from a product development and model development perspective and as part of the culture and hopefully that's valuable for others as well.

Highlight slides
Core thesis: frontier models + frontier products✦ from: Frontier models and frontier products create a mutually reinforcing flywheel — Opus 45 would not have had its moment without Claude Code, and Claude Code would not have seen accelerated adoption without Opus 45.Two inflection points, one flywheel✦ from: Frontier models and frontier products create a mutually reinforcing flywheel — Opus 45 would not have had its moment without Claude Code, and Claude Code would not have seen accelerated adoption without Opus 45.Scaling loss is smooth; capabilities jump✦ from: AI capabilities emerge in discontinuous jumps rather than smooth improvements — models can go from unable to do something to reliably doing it, which makes safety harder and demands that organizations have eval systems in place to detect these jumps when they happen.Discontinuous jumps demand eval systems✦ from: AI capabilities emerge in discontinuous jumps rather than smooth improvements — models can go from unable to do something to reliably doing it, which makes safety harder and demands that organizations have eval systems in place to detect these jumps when they happen.Evals are the new PRDs✦ from: For research product managers at Anthropic, evals are the new PRDs — the primary artifact for defining user value has shifted from documents describing what to build to reproducible tests that measure whether the model has improved on specific user pain points.Evals personify user needs✦ from: For research product managers at Anthropic, evals are the new PRDs — the primary artifact for defining user value has shifted from documents describing what to build to reproducible tests that measure whether the model has improved on specific user pain points.
Related episodes