ATRIUMsearch → argument graph
MechanismArticle

During DeepSeek-R1's RLVR training, the reasoning trace itself was not used as a training signal — the reward was based only on the final answer's correctness and format, since including the trace in training was reported to not help.

Raschka explains that RLVR rewards only the final answer and format, not the intermediate reasoning trace, per the DeepSeek-R1 paper's findings. ✦ AI generated

Sebastian Raschka · Ahead of AI · 2026-07-18 · original ↗

Notably, the reasoning trace itself was not used for training or updating the model. Although they tried to use this intermediate response information for training, the DeepSeek-R1 paper reported that it wasn't helpful for the model training, so it was ultimately not used.

Read full article ↗excerpt · fair-use quotation

Around this claim