DataArticle
Kimi K3 hit 1668 Elo on GDPval v2, 53% and #1 on AutomationBench-AA, and 1547 Elo on AA-Briefcase, at $0.94 cost per task and about 21% fewer output tokens than K2.6 across the full Intelligence Index run.
Artificial Analysis published detailed independent benchmark numbers showing K3 leading AutomationBench-AA and using notably fewer output tokens than its predecessor K2.6. ✦ AI generated
Artificial Analysis · Latent Space · 2026-07-17 · original ↗
AA also reported K3 at 1668 Elo on GDPval v2, 53% / #1 on AutomationBench-AA, and 1547 Elo on AA-Briefcase, with cost per task of $0.94, about 21% fewer output tokens than K2.6 across the full Intelligence Index run
Read full article ↗excerpt · fair-use quotation
Around this claim
Counterpoint · 2
Kimi K3's benchmark story might be overstated unless validated on hidden or uncontaminated evals like LiveBench, and if the model 'thinks forever,' its real-world cost could end up less favorable than advertised.Bindu Reddy · Latent Space · conf 70%Kimi K3 currently offers only one reasoning effort level, 'max,' which makes it expensive to run: it burned 13,241 reasoning tokens to produce just 3,417 tokens of output, costing 25 cents for a single pelican SVG.Simon Willison · Simon Willison's Weblog · conf 60%