Mechanism◆Video · 20:36 · 2m
Standard RL training rewards an entire trajectory as pass/fail without indicating which specific decisions were good or bad, which is like a teacher grading an essay 'B' without saying what made it a B — attributing credit to the specific pivotal decisions within a trajectory would make RL training far more efficient.
Alistair Pullen · Machine Learning Street Talk