Mechanism◆Article
The infrastructure reality is that 'open-weight' does not mean easy to run: a 2.4T-class MoE like Qwen3.8-Max is not a commodity local model, requiring massive memory and accelerator counts similar to Kimi K3.
Because a 2.4T-class MoE needs over 1TB of memory and at least 8 H100/B200 GPUs (with Moonshot recommending 64+ accelerators for K3), 'open-weight' is operational openness, not local-inference accessibility — a key split between flagship and 27B releases. ✦ AI generated
AINews · Latent Space · 2026-08-04 · original ↗
Jamin Ball argued that pricing comparisons were overstated because 'vanilla' token prices ignore token efficiency and because these models are enormous: Qwen 3.8 Max >2T params; Kimi K3 ~104B active per token; GLM 5.2 = 744B total, 40B active. For K3, loading weights alone is >1TB memory. Requires at least 8 H100/B200 GPUs to run. Moonshot recommends 64+ accelerators in supernode-style setups. This same critique implicitly applies to Qwen3.8-Max, even if its active-parameter count is somewhat lower than K3's: a 2.4T-class MoE is not a commodity local model. StableQuan made the practical version of the same point more bluntly: long, RAM-heavy prompts and slow tool calls make giant models painful on consumer hardware, recommending API use instead.
Read full article ↗excerpt · fair-use quotation
- ·Large MoE models need massive infrastructure; not commodity local models
- ·Qwen3.8-Max and Kimi K3 are operational-open, not local-inference accessible
- ·Loading K3 weights alone exceeds 1TB memory
- ·Slit between flagship and 27B-class local releases
- ·K3: >1TB memory to load weights, 8+ H100/B200 GPUs minimum
- ·Moonshot recommends 64+ accelerators in supernode setups
- ·Same critique applies to Qwen3.8-Max, a 2.4T-class MoE
- ·Giant models painful on consumer hardware; API use advised
Around this claim