ATRIUMsearch → argument graph
Article · 2026-07-17 · 6 moments

[AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing

a great week for open models continues. ✦ AI generated

01
Claim

Despite being highly competitive overall, Kimi K3 still has a noticeable gap in user experience versus Claude Fable 5 and GPT-5.6 Sol.

Moonshot itself conceded that while K3 is highly competitive on capability, it still trails top closed models in overall user experience.

transcript

Moonshot AI: Moonshot's own phrasing acknowledged a limitation: despite being highly competitive overall, K3 still has a "noticeable gap in user experience" versus Claude Fable 5 and GPT-5.6 Sol

extends · 1provides context · 1supports · 1

02
Data

Kimi K3 became #1 in Frontend Code Arena with 1679 points, surpassing Claude Fable 5 and jumping from #18 (as K2.6) to #1, ranking #1 in 6 of 7 frontend domains and #2 in Gaming.

Arena's pairwise human-preference leaderboard put Kimi K3 in first place for frontend code generation, ahead of Claude Fable 5, a major jump from its predecessor's rank.

transcript

Arena: Arena then reported a major early result: Kimi K3 became #1 in Frontend Code Arena with 1679 points, surpassing Claude Fable 5 and jumping from #18 (K2.6) to #1, ranking #1 in 6 of 7 frontend domains and #2 in Gaming

03
Data

Kimi K3 hit 1668 Elo on GDPval v2, 53% and #1 on AutomationBench-AA, and 1547 Elo on AA-Briefcase, at $0.94 cost per task and about 21% fewer output tokens than K2.6 across the full Intelligence Index run.

Artificial Analysis published detailed independent benchmark numbers showing K3 leading AutomationBench-AA and using notably fewer output tokens than its predecessor K2.6.

transcript

Artificial Analysis: AA also reported K3 at 1668 Elo on GDPval v2, 53% / #1 on AutomationBench-AA, and 1547 Elo on AA-Briefcase, with cost per task of $0.94, about 21% fewer output tokens than K2.6 across the full Intelligence Index run

rebuts · 2

04
Context

Kimi K3 is an 'Open Frontier Intelligence' model with 2.8T total parameters, 1M-token context, native multimodal input, Kimi Delta Attention (KDA), and Attention Residuals, live on Kimi.com, Kimi Work, Kimi Code, and API, with open weights promised by July 27, 2026.

Moonshot AI officially unveiled Kimi K3, a 2.8T-parameter open-weights model with a 1M-token context window and novel attention architecture, calling it 'Open Frontier Intelligence.'

transcript

Moonshot AI: Moonshot officially introduced Kimi K3 as "Open Frontier Intelligence" with 2.8T total parameters, 1M-token context, native multimodal input, Kimi Delta Attention (KDA), and Attention Residuals, and said the model is live on Kimi.com, Kimi Work, Kimi Code, and API, with open weights promised by July 27, 2026

extends · 1provides context · 1

05
Prediction

Kimi K3's benchmark story might be overstated unless validated on hidden or uncontaminated evals like LiveBench, and if the model 'thinks forever,' its real-world cost could end up less favorable than advertised.

Bindu Reddy cautioned that K3's impressive leaderboard results could be inflated by benchmark contamination and that its verbose 'thinking' behavior might erode its apparent cost advantage in practice.

transcript

Bindu Reddy: Bindu Reddy warned that K3's benchmark story might be overstated unless validated on hidden / uncontaminated evals like LiveBench, and argued that if the model "thinks forever," real cost could be less favorable

explains mechanism · 1supports · 1

06
Data

Kimi K3 scores 57 on the AA Intelligence Index, comparable to Opus 4.8 and GPT-5.5, but still behind Fable 5 and GPT-5.6 Sol overall.

Artificial Analysis' independent index placed K3 in the tier of Opus 4.8 and GPT-5.5, short of the very top closed models, offering a more measured read than the headline hype.

transcript

Artificial Analysis: Artificial Analysis published an independent evaluation placing K3 at 57 on the AA Intelligence Index, calling it comparable to Opus 4.8 and GPT-5.5, but still behind Fable 5 and GPT-5.6 Sol overall

provides context · 2

Highlight slides
Related episodes