Mechanism◆Audio · 60:33 · 2m
MCTS-based training works because it never asks the network to directly chase a win/loss signal — instead, for every single action taken, search produces a strictly better relabeled target, giving one clean supervision signal per action instead of a noisy win-rate credit-assignment problem.
Eric Jang · Dwarkesh Podcast