DataArticle
The third-party eval and leaderboard results place Qwen3.8-Max in the same deployment class as other giant sparse open models like Kimi K3 and GLM-5.2, not the practical 30B–70B local tier.
Third-party results are strong and consistent: 1M context, 128k output, #4 in Frontend Code Arena, #2 in Vision Arena, 66.1 on Vals Index with a +8.6 gain in ~2.5 months, 87.3% SWE-bench, and 2.3x lower cost-per-test than Claude Opus 4.7 — positioning it alongside Kimi K3 and GLM-5.2 rather than local models. ✦ AI generated
AINews · Latent Space · 2026-08-04 · original ↗
Frontend Code Arena: Qwen3.8-Max debuted at #4 overall with 1,668 Elo, trailing only Claude Opus 5 [Max] at 1,705 and Kimi K3 [Max] at 1,676... Vision Arena: Qwen3.8-Max ranked #2 with 1,305, only 13 points behind Claude Fable 5 [High]. Vals Index: Qwen3.8-Max ranked #2 among open-weight models, #10 overall out of 43, with a score of 66.1. It matched Claude Opus 4.7 on the Index, 66.1 vs 66.1. At about 2.3x lower cost per test: $2.68 vs $6.17. SWE-bench: 87.3%, ahead of GPT-5.5 (82.6%) and GLM-5.2 (83.3%), but behind Claude Opus 4.8 (89.2%). Vals also highlighted the pace of progress: Qwen 3.7 Max = 57.5, Qwen 3.8 Max = 66.1, gain of 8.6 points in ~2.5 months. These numbers matter because they place Qwen3.8-Max in the same deployment class as other giant sparse open models like Kimi K3 and GLM-5.2, not the more practical 30B–70B local tier.
Read full article ↗excerpt · fair-use quotation
Around this claim