MechanismArticle
Frontier models perform better and produce broader coverage when given room to decide how to approach a task rather than being handed a detailed checklist of steps.
Claire found that switching from listing 25 specific testing steps to simply saying 'QA the onboarding flow' produced broader coverage with fewer blind spots introduced by her own assumptions. ✦ AI generated
Claire · Lenny's Newsletter · 2026-07-27 · original ↗
Frontier models often perform better when they are given room to think. When Claire first started using browser use, she would give the model a list of 25 things to test. Now she simply says, 'QA the onboarding flow,' and lets it decide how to approach the task. The result is often broader coverage, with fewer blind spots introduced by her own assumptions about what matters.
Read full article ↗excerpt · fair-use quotation
Around this claim
In practice · 2
AI testing can be far more exhaustive than human testing because AI does not default to following the happy path, as demonstrated when Codex immediately found a blocking bug in an onboarding flow that had survived for months.Claire · Lenny's Newsletter · conf 75%Persona testing with browser use — having AI use a product as a specific persona like a PM coming out of a meeting — can reveal structural friction that synthetic user research misses.Claire · Lenny's Newsletter · conf 65%