ATRIUMsearch → argument graph
DataArticle

FTPO (Final Token Preference Optimization) training, branded Antidoom, sharply cuts doom-loop rates in small reasoning models: LFM2.5-2.6B falls from 10.2% to 1.4% and Qwen3.5-4B falls from 22.9% to 1% under greedy sampling, alongside downstream eval gains.

Liquid AI's open-source Antidoom method relabels the token that triggers repetitive 'doom loops' and redistributes probability toward alternatives, yielding large measured drops in loop rates for LFM2.5-2.6B and Qwen3.5-4B. ✦ AI generated

Liquid AI · Latent Space · 2026-07-08 · original ↗

The reported reductions are substantial: LFM2.5-2.6B from 10.2% → 1.4% and Qwen3.5-4B from 22.9% → 1% under greedy sampling, with downstream eval gains. The method, FTPO (Final Token Preference Optimization), relabels the loop-triggering token and redistributes probability toward alternatives.

Read full article ↗excerpt · fair-use quotation

Related moments