ATRIUMsearch → argument graph
MechanismArticle

The infrastructure reality is that 'open-weight' does not mean easy to run: a 2.4T-class MoE like Qwen3.8-Max is not a commodity local model, requiring massive memory and accelerator counts similar to Kimi K3.

Because a 2.4T-class MoE needs over 1TB of memory and at least 8 H100/B200 GPUs (with Moonshot recommending 64+ accelerators for K3), 'open-weight' is operational openness, not local-inference accessibility — a key split between flagship and 27B releases. ✦ AI generated

AINews · Latent Space · 2026-08-04 · original ↗

Jamin Ball argued that pricing comparisons were overstated because 'vanilla' token prices ignore token efficiency and because these models are enormous: Qwen 3.8 Max >2T params; Kimi K3 ~104B active per token; GLM 5.2 = 744B total, 40B active. For K3, loading weights alone is >1TB memory. Requires at least 8 H100/B200 GPUs to run. Moonshot recommends 64+ accelerators in supernode-style setups. This same critique implicitly applies to Qwen3.8-Max, even if its active-parameter count is somewhat lower than K3's: a 2.4T-class MoE is not a commodity local model. StableQuan made the practical version of the same point more bluntly: long, RAM-heavy prompts and slow tool calls make giant models painful on consumer hardware, recommending API use instead.

Read full article ↗excerpt · fair-use quotation

Around this claim