ATRIUMsearch → argument graph
MechanismVideo · 40:19 — 46:25

For research product managers at Anthropic, evals are the new PRDs — the primary artifact for defining user value has shifted from documents describing what to build to reproducible tests that measure whether the model has improved on specific user pain points.

Dianne describes how the PM role has fundamentally changed: instead of writing PRDs, AI product managers translate user feedback into evals — structured tests that capture the exact failure trajectory (hallucination, overconfidence, tool-calling error, etc.) so researchers can measure improvement across model generations. ✦ AI generated

Dianne Penn · Lenny's Podcast · 2026-07-26 · original ↗

starts at this moment · 40:19

Elicited by

When you're hiring PMs, product people, when you're looking at people that do well in today's world, what are some things that you notice? What are you looking for more most? What's kind like trending up in what you find is important and what's kind of trending down?

I think you think of a product manager as I own product strategy and delivering user value as but I demonstrate day-to-day by writing a PRD or writing a product vision doc. And for for my team as like research product managers, the way to drive user value is to figure out the right user feedback, the evals, right? That then can be a personification of that user need. So like we do write some product documents and PRDs, but we actually have a saying on the team of evals are the new PRDs, right? because in order to deliver that user value uh it's not that exact artifact that people used to write in the last like one to two decades it's a new way of working and so the first think principal thinking would be let me figure out what is the thing I should do to achieve my goals rather than here is a set of activities that I've done and therefore I will continue to do.

verbatim transcript · starts at 40:19

Transcript · around this moment

40:19generalists like PM's generalist like research product managers have actually been the same. Um so I think some of those traits number one is first principles thinking and this is really uh rather than pattern matching what you used to do in let's say consumer product or B2B SAS um but actually figuring out in this moment for this user group with this technology what what is the user value

40:50>> is there an example that a lot of people hear first principles thinking they're like yes I about it. I'm good at this. What is what's an example of someone having really demonstrated really good first principles thinking? >> I think one example is I think you think of a product manager as I own product strategy and delivering user value as but I demonstrate day-to-day by writing a PRD or writing a product vision doc.

41:16And for for my team as like research product managers, the way to drive user value is to figure out the right user feedback, the evals, right? That then can be a personification of that user need. So like we do write some product documents and PRDs, but we actually have a saying on the team of evals are the new PRDs, right? because in order to deliver that user value uh it's not that exact

41:49artifact that people used to write in the last like one to two decades it's a new way of working and so the first think principal thinking would be let me figure out what is the thing I should do to achieve my goals rather than here is a set of activities that I've done and therefore I will continue to do >> so the idea here is used to be have kind

42:12of an idea create a PRD talk to people about it. Align on the plan, design it, build it, ship it, see how it goes, iterate. What I'm hearing here is it's like, okay, here's some feedback about something that's wrong or an opportunity. Step one is the eval is now how you define what the work is versus a PRD. >> Maybe maybe step one would be uh understanding the user painoint. And so

42:39the way to even access that user painpoint is different, right? In the past, we might do a user interview and I think if you go like deep enough, you you might have the user walk you through their user flow, the pixels. Here, you have to sweat the tokens as much as you sweat the pixels. And so, one activity we have on the team is reading the transcripts and understanding

43:04uh what was the trajectories that failed very deeply to then say was this like a hallucination? was this claw being overconfident. So like the theme of the failure actually has a lot of nuance and then that allows you to build a description a like sustained description of that painoint. Uh so that could be essentially in a new eval and is the eval on distribution right is it capturing both the positive

43:37situations where this is failing and also areas when it should actually not fail and then bring that back to let's say research so then we can make the improvements and actually measure the quality of okay when we have opus 5.5 is this area improving or not is claude now able to uh identify the right places in the document uh and pull the right synthesis out. So it's just the

Around this claim