ATRIUMsearch → argument graph
ClaimArticle · 33:20 — 34:50

Distillation is becoming less impactful over time as training shifts to RL, and Ben Thompson's contrary claim that distillation helps more during RL is wrong.

Nathan argues that distillation via SFT was more impactful in earlier generations, but now that RL dominates post-training, the marginal benefit of distilling from frontier models has diminished. He pushes back on Ben Thompson's claim that RL makes distillation more important. ✦ AI generated

Nathan Lambert · Interconnects · 2026-07-22 · original ↗

Distillation has become less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to RL. Ben Thompson very strongly proclaimed that distillation is getting more impactful as you do RL. But distillation during the RL stage is a lot harder. Big RL runs are millions of rollouts — to do this on an API like Fable would be insanely expensive and might not even give you a performance uplift versus using your own tailored grader model.

Read full article ↗excerpt · fair-use quotation

Around this claim