ClaimArticle
The book communicates the intuitions and history behind post-training, explaining in simple terms why post-training works, the trade-offs to get it right, and the misconceptions people get stuck on.
Most of the book is about communicating intuitions and history of post-training, aimed at explaining why post-training works, the trade-offs needed, and the misconceptions people fall into, with an intentionally higher-voice explanatory style than typical textbooks. ✦ AI generated
Nathan Lambert · Interconnects · 2026-08-10 · original ↗
Otherwise, most of the book is about communicating intuitions and history. Much of the LLM industry is defined by core techniques that haven't changed much in the last few years. This book was my attempt to explain in simple terms why post-training works, what trade-offs people need to make to get it right, and what misconceptions people often get stuck on.
Read full article ↗excerpt · fair-use quotation
Around this claim
This moment responds to
explains mechanism → My post-training textbook 'Reinforcement Learning from Human Feedback: Aligning and Post-training LLMs' is complete and now shipping via Manning, Amazon US, and (in October) Amazon UK.Nathan Lambert · Interconnectsexplains mechanism → The book existed because critical post-training methods, such as rejection sampling, outcome reward models, and character training, had no foundational online material explaining them and still lack such resources today.Nathan Lambert · Interconnectsextends → Post-training-based RL intuition, from the policy-gradient theorem through PPO to modern versions like GSPO and CISPO, is essential for judging whether a new algorithm is fake or has real potential.Nathan Lambert · Interconnects