ClaimArticle
Vendor benchmark claims are unreliable—scorer bugs can move a system's reported score from 65% to 93.6%—and developers should run their own evaluations rather than trusting marketing claims, including those from their own organization.
A recurring theme in the AI community is skepticism toward vendor benchmarks, highlighted by concrete examples like scorer bugs creating massive score swings and François Chollet's reminder that ARC-3 leaderboard scores are weak proxies for real-world performance. ✦ AI generated
AINews · Latent Space · 2026-08-14 · original ↗
The eval backlash continues: A recurring theme was skepticism toward vendor benchmark claims. Vik Paruchuri criticized a LlamaIndex benchmark, saying scorer bugs could move a system from 65% to 93.6%, and explicitly argued developers should run their own evals rather than trust marketing—'including ours'. François Chollet reiterated that the public ARC-3 demonstration set is not training or eval data and that leaderboard scores there are weak proxies for private-set performance.
Read full article ↗excerpt · fair-use quotation
Around this claim