Mechanism◆Audio · 36:10 — 37:40
Smaller models can be far more capable than expected if trained for persistence, verification, and backtracking behaviors rather than raw intelligence alone.
Eiso explains that Laguna S (118B total, 8B active) outperforms models many times its size because post-training instilled behaviors like persistence and backtracking, not just raw intelligence. He suggests this means the peak of model usefulness for knowledge work may be much closer than previously thought. ✦ AI generated
Eiso Kant · Latent Space · 2026-07-23 · original ↗
plays this moment only · 36:10 — 37:40
The gains in Laguna S come not from more intelligence, but more from different behavior, more verification, less taking things for granted, not declaring victory early, and being way more persistent... This model for me is the first sign that maybe that peak is at a trillion, five trillion, ten trillion. Maybe we can just squeeze way more out of these models.
verbatim transcript · starts at 36:10
Transcript · around this moment
36:07Laguna S: Persistence vs. Raw Intelligence
- ·118B total / 8B active parameters outperforms much larger models
- ·Gains come from behavior (persistence, verification, backtracking), not raw intelligence
- ·Post-training instilled: not declaring victory early, not taking things for granted
- ·First sign that the capability peak may be at 1T–10T parameters
- ·Suggests we can squeeze far more out of existing small models
- ·Implies diminishing returns from scale alone if behavior isn't trained
Around this claim
This moment responds to
explains mechanism → Onyx trains small, specialized models that act as a triage layer—capable only of deciding whether a more expensive, smarter agent needs to intervene—solving the cost and latency problem that would otherwise make agent oversight economically impossible.Maxim Bar Kogan · No Priorssupports → Model size and reasoning effort are two separate scaling knobs whose performance curves overlap — a smaller model run at higher reasoning effort can sometimes match a larger model run at lower effort.Sebastian Raschka · Ahead of AIgives example → A base model's accuracy on a math benchmark can jump from 15% to 50% after only 50 steps of RLVR training, showing that RL is unlocking latent pre-trained knowledge rather than teaching genuinely new mathematical understanding.Sebastian Raschka · Lex Fridman