DataArticle
On the Harvey LAB-AA legal-agent benchmark, models can satisfy many individual rubric items yet still fail to produce acceptable end-to-end deliverables, with even the leading model, Claude Fable 5, achieving only a 14.2% all-pass rate.
Artificial Analysis's new Harvey LAB-AA benchmark, covering 120 private legal tasks across 24 practice areas, shows top models like Claude Fable 5 leading at only a 14.2% all-pass rate, underscoring a gap between partial-credit performance and full task success. ✦ AI generated
Artificial Analysis · Latent Space · 2026-07-08 · original ↗
Artificial Analysis launched Harvey LAB-AA, a legal-agent benchmark over 120 private legal tasks across 24 practice areas, where Claude Fable 5 led at 14.2% all-pass rate; Claude Opus 4.8 and GLM-5.2 tied at 7.5%, with GLM hitting that at roughly ~6% of Fable's cost per task in their release.
Read full article ↗excerpt · fair-use quotation
Around this claim
This moment responds to
gives example → Benchmarks that only report the percentage of tests an AI-generated program passes are misleading, because a program that passes just 70-80% of tests is probably not actually correct — what matters is whether it got any test fully right, not the aggregate score.Thomas Ahle · Machine Learning Street Talkrebuts → The billable hour business model at law firms only survives by overcharging junior associates' hourly rates to compensate for underpricing the much higher value of senior partners' time, and AI is starting to break that model.Max · All-In Podcast