ATRIUMsearch → argument graph
DataArticle

On the Harvey LAB-AA legal-agent benchmark, models can satisfy many individual rubric items yet still fail to produce acceptable end-to-end deliverables, with even the leading model, Claude Fable 5, achieving only a 14.2% all-pass rate.

Artificial Analysis's new Harvey LAB-AA benchmark, covering 120 private legal tasks across 24 practice areas, shows top models like Claude Fable 5 leading at only a 14.2% all-pass rate, underscoring a gap between partial-credit performance and full task success. ✦ AI generated

Artificial Analysis · Latent Space · 2026-07-08 · original ↗

Artificial Analysis launched Harvey LAB-AA, a legal-agent benchmark over 120 private legal tasks across 24 practice areas, where Claude Fable 5 led at 14.2% all-pass rate; Claude Opus 4.8 and GLM-5.2 tied at 7.5%, with GLM hitting that at roughly ~6% of Fable's cost per task in their release.

Read full article ↗excerpt · fair-use quotation

Around this claim