ATRIUMsearch → argument graph
Article · 2026-07-16 · 6 moments

[AINews] Thinky's Inkling: 975B-A41B multimodal, new best American Apache 2.0 open model (with Inkling-Small, 276B-A12B)

Thinky's first full LLM release is a banger and bonus: it's open weights! ✦ AI generated

01
Data

Inkling debuts at 41 on the Intelligence Index, making it the leading U.S. open-weights release, ahead of Nemotron 3 Ultra (38), Gemma 4 31B (29), and gpt-oss-120b (24).

Artificial Analysis's independent benchmark places Inkling at the top of U.S. open-weight models on its Intelligence Index, ahead of Nemotron 3 Ultra, Gemma 4 31B, and gpt-oss-120b.

transcript

Artificial Analysis: Artificial Analysis said Inkling debuts at 41 on the Intelligence Index, making it the leading U.S. open-weights release and ahead of Nemotron 3 Ultra (38), Gemma 4 31B (29), and gpt-oss-120b (24).

02
Claim

Inkling does in fact use distillation from open weights and underperforms GLM 5.2 on TerminalBench-style evals, contradicting hype that it was trained entirely without distillation.

JJitsev disputes the narrative that Inkling is uniquely 'pure' among open-weight models, pointing out it does use distillation and still lags GLM 5.2 on terminal/agentic benchmarks.

transcript

JJitsev: JJitsev pushed back on hype around 'only open-weight model trained without distilling,' saying Inkling uses distillation from open weights and underperforms GLM 5.2 on TerminalBench-style evals.

03
Claim

Inkling is a clear step up from Nemotron Ultra and the new best American open model, though still a bit behind GLM 5.2 on agentic benchmarks and Kimi K2.6 on multimodal tasks.

Nathan Lambert judges Inkling as the leading U.S. open model to date, while noting it still trails top Chinese open models like GLM 5.2 and Kimi K2.6 on specific capabilities.

transcript

Nathan Lambert: Natolambert called it a 'clear step up from Nemotron Ultra' and 'new best American model,' but still 'a bit behind GLM 5.2 on agentic benchies, and Kimi K 2.6 on multi modal'.

supports · 1

04
Fact

Inkling is Thinking Machines Lab's first general model: fully open-weight, 975B total parameters, natively multimodal, and available on Tinker, Hugging Face, and partner platforms.

Soumith Chintala describes Inkling as Thinking Machines' first general-purpose model, emphasizing its open weights, huge parameter count, and native multimodal support across major platforms.

transcript

Soumith Chintala: Soumith Chintala framed it as Thinking Machines' 'first general model,' stressing open weights, 975B parameters, native multimodality, and availability on Tinker, Hugging Face, and partners.

extends · 4supports · 1

05
Claim

Inkling's benchmark results are not that great, effectively amounting to 'another Kimi-K2.6', behind all closed models and GLM-5.2, and possibly rushed out ahead of Kimi-K3 and DeepSeek-V4-GA.

Scaling01 pushes back on the hype, arguing Inkling's benchmarks are mediocre and comparable to older Chinese open models, speculating the release timing was defensive.

transcript

Scaling01: Scaling01 argued the benchmarks are 'not that great,' describing it as roughly 'another Kimi-K2.6' and behind all closed models and GLM-5.2, speculating the release may have been timed ahead of Kimi-K3 and DeepSeek-V4-GA.

rebuts · 1supports · 1

06
Mechanism

Inkling uses relative positional encoding / relative attention bias instead of RoPE, which multiple technical observers called one of the most novel large-scale architecture choices in the release.

Community technical analysts highlight that Inkling forgoes the now-standard RoPE positional encoding in favor of relative attention bias, seen as a bold, closely-watched architecture bet at this scale.

transcript

eliebakouch: Relative positional encoding / relative attention bias instead of RoPE; multiple posters called this one of the most novel large-scale choices.

supports · 1

Highlight slides
Related episodes