ATRIUMsearch → argument graph
DataArticle

Nanbeige4.2-3B, a 3B model using a Looped Transformer that reuses layers, outperforms models 4x its size on agent and reasoning benchmarks.

A new 3B agentic model using a Looped Transformer with layer reuse reportedly outperforms Qwen3.5-9B and Gemma4-12B on several benchmarks, with commenters noting the architectural implications for parameter efficiency. ✦ AI generated

Reddit (r/LocalLlama post) · Latent Space · 2026-07-23 · original ↗

The image is a technical benchmark bar chart supporting the post's claim that Nanbeige4.2-3B, a 3B non-embedding-parameter agentic model using a Looped Transformer that reuses layers, can outperform larger models such as Qwen3.5-9B and Gemma4-12B on several agent/reasoning/code benchmarks.

Read full article ↗excerpt · fair-use quotation

Related moments