ClaimArticle
Kimi K3's strong benchmark scores are likely inflated by benchmark-specific tuning that doesn't reflect true generalization.
Jack Clark notes Kimi K3 scores near frontier Western models on benchmarks but suspects 'benchmaxxing' — tuning aimed at benchmarks rather than real generalization. ✦ AI generated
Jack Clark · Import AI · 2026-07-20 · original ↗
However, Kimi has some brittleness which smells to me like "benchmaxxing" - performance may have been tuned around these benchmarks in a way that harms some parts of generalization.
Read full article ↗excerpt · fair-use quotation
Around this claim