ATRIUMsearch → argument graph
Article · 2026-08-04 · 6 moments

[AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork

Qwen is so back! ✦ AI generated

01
Data

Qwen3.8-Max demonstrates autonomous long-horizon and multimodal capabilities including 10+ days of unattended coding, a 125-hour autonomous research loop that beat an original paper's benchmark by +2.71 points, silicon design from RTL to physical layout, and product-grade outputs across professional workflows.

Alibaba frames Qwen3.8-Max around headline autonomous capabilities: building a self-evolving coding harness over 10+ days, inventing a data-selection method beating a paper by +2.71 points across 125 hours, completing a full silicon design flow, and beating 87% of human teams in a data science challenge.

transcript

AINews: 10+ Days Unattended Coding: Built a self-evolving coding harness from scratch over a multi-week autonomous run. Autonomous AI Research: Rebuilt a complete paper's pipeline (Unified Data Selection for LLM Reasoning) from scratch, then autonomously ran an iterative research loop over 125 hours to invent a new data selection method beating the original paper's benchmark by +2.71 points. Competitive Data Science: Competed against 526 human teams in the WWW2025 Multimodal Dialogue Intent Recognition Challenge, placing in the top 13% (outperforming 87% of human teams) within 24 hours. Executed a complete silicon design flow (GCD/RSA cryptographic accelerator) from RTL editing to simulation, synthesis, and physical layout. Reduced gate count from 8,298 to 678 gates while achieving an 81% die area reduction and meeting physical timing closure at 500 MHz. Demonstrated production-grade outputs across hundreds of professional workflows (e.g., corporate legal reviews, UI/UX design, structural engineering models, and automated ETF quant research). Outperformed competing models in the E-Commerce Bench (a 365-day store operation simulation), generating a 4.16x return (¥416,252 balance) through continuous game-theoretic negotiation and inventory planning. Integrates native visual feedback across planning, coding, and GUI interaction, enabling direct application recreation across platforms (desktop, mobile, web).

provides context · 2rebuts · 1supports · 1

02
Claim

Qwen3.8-Max is a 2.4T-parameter flagship model that would have been the top open model in the world but for the Kimi K3 release, and Alibaba has promised to open-weight both it and the 27B variant.

After a year of doubt over whether Alibaba would keep releasing relevant open models, Qwen is 'so back' with a monster 2.4T flagship, Qwen3.8-Max, priced at $2/$6 input/output, with open weights promised for both it and a 27B model.

transcript

AINews: After the Qwen Exodus last year and new management took over launching more closed model APIs, there was some real doubt as to whether or not this leading open models lab would continue to release relevant models. That doubt is now gone. Qwen 3.8 Max is a MONSTER 2.4T model that would have been the top open model in the world but for the Kimi K3 release we already covered. Qwen offers them on API for $2 input/$6 output per million tokens, but they have promised to open-weight both models.

explains mechanism · 1extends · 1provides context · 1

03
Prediction

Qwen3.8-Max represents a strategic shift by Alibaba toward ecosystem influence over exclusivity and the shrinking of the moat to post-training, agent harnesses, inference infra, and developer lock-in rather than raw pretraining.

ZhihuFrontier frames the release as Alibaba choosing ecosystem adoption over keeping top-tier systems API-only, pushed by DeepSeek and Kimi eroding the premium on closed APIs; the same release feeds a broader narrative of China dominating the open-weight frontier while US labs lead mainly in closed models.

transcript

AINews: ZhihuFrontier explicitly framed the move as Alibaba choosing ecosystem influence over exclusivity, arguing that earlier Max models stayed closed while the open line had previously topped out around Qwen3-235B. In that reading, DeepSeek, Kimi, and other Chinese open models weakened the premium of keeping top-tier systems API-only, pushing Alibaba to compete on ecosystem adoption as well as model quality... A subtext here is that the moat may be shifting: not just raw pretraining, but post-training, agent harnesses, inference infra, distillation pipelines, and developer lock-in. That is exactly why an open-weight flagship at 2.4T is strategically valuable even if relatively few teams ever self-host it.

provides context · 1

04
Mechanism

The infrastructure reality is that 'open-weight' does not mean easy to run: a 2.4T-class MoE like Qwen3.8-Max is not a commodity local model, requiring massive memory and accelerator counts similar to Kimi K3.

Because a 2.4T-class MoE needs over 1TB of memory and at least 8 H100/B200 GPUs (with Moonshot recommending 64+ accelerators for K3), 'open-weight' is operational openness, not local-inference accessibility — a key split between flagship and 27B releases.

transcript

AINews: Jamin Ball argued that pricing comparisons were overstated because 'vanilla' token prices ignore token efficiency and because these models are enormous: Qwen 3.8 Max >2T params; Kimi K3 ~104B active per token; GLM 5.2 = 744B total, 40B active. For K3, loading weights alone is >1TB memory. Requires at least 8 H100/B200 GPUs to run. Moonshot recommends 64+ accelerators in supernode-style setups. This same critique implicitly applies to Qwen3.8-Max, even if its active-parameter count is somewhat lower than K3's: a 2.4T-class MoE is not a commodity local model. StableQuan made the practical version of the same point more bluntly: long, RAM-heavy prompts and slow tool calls make giant models painful on consumer hardware, recommending API use instead.

provides context · 1

05
Context

The licensing ambiguity is the most concrete skeptical reaction: a flagged license prohibition covering the USA, EU, UK, and Korea would mean 'open weights' lacks OSI-style rights, imposes use-case and jurisdiction restrictions, and may forbid even downloading the model from the US.

OstrisAI read the Qwen license as prohibiting use in the USA, EU, UK, and Korea, echoing the simultaneous MiniMax H3 licensing debate; no clarifying tweet from Alibaba resolved the restrictive-license reading, making licensing ambiguity a main reason reception was cautious.

transcript

AINews: OstrisAI flagged what they read as a license prohibition covering the USA, EU, UK, and Korea, saying the terms appeared to forbid even downloading the model from the US... That concern echoed a broader discussion happening simultaneously around another open-weight release, MiniMax H3, where users argued that geographic restrictions undercut claims of openness. No clarifying Qwen license tweet appears in this dataset from Alibaba itself, so the restrictive-license reading remained unresolved within these tweets. For engineers, this matters more than the marketing label. 'Open weights' can still mean: no OSI-style open-source rights, use-case restrictions, export/jurisdiction limits, or no legal permission for commercial deployment in key regions. That licensing ambiguity is one of the main reasons some of the reaction was more cautious than celebratory.

explains mechanism · 1provides context · 2

06
Data

The third-party eval and leaderboard results place Qwen3.8-Max in the same deployment class as other giant sparse open models like Kimi K3 and GLM-5.2, not the practical 30B–70B local tier.

Third-party results are strong and consistent: 1M context, 128k output, #4 in Frontend Code Arena, #2 in Vision Arena, 66.1 on Vals Index with a +8.6 gain in ~2.5 months, 87.3% SWE-bench, and 2.3x lower cost-per-test than Claude Opus 4.7 — positioning it alongside Kimi K3 and GLM-5.2 rather than local models.

transcript

AINews: Frontend Code Arena: Qwen3.8-Max debuted at #4 overall with 1,668 Elo, trailing only Claude Opus 5 [Max] at 1,705 and Kimi K3 [Max] at 1,676... Vision Arena: Qwen3.8-Max ranked #2 with 1,305, only 13 points behind Claude Fable 5 [High]. Vals Index: Qwen3.8-Max ranked #2 among open-weight models, #10 overall out of 43, with a score of 66.1. It matched Claude Opus 4.7 on the Index, 66.1 vs 66.1. At about 2.3x lower cost per test: $2.68 vs $6.17. SWE-bench: 87.3%, ahead of GPT-5.5 (82.6%) and GLM-5.2 (83.3%), but behind Claude Opus 4.8 (89.2%). Vals also highlighted the pace of progress: Qwen 3.7 Max = 57.5, Qwen 3.8 Max = 66.1, gain of 8.6 points in ~2.5 months. These numbers matter because they place Qwen3.8-Max in the same deployment class as other giant sparse open models like Kimi K3 and GLM-5.2, not the more practical 30B–70B local tier.

provides context · 1

Highlight slides
Related episodes