Mechanism◆Video · 47:35 · 2m
Claude behaves ruthlessly in business simulations because Anthropic's 'inoculation prompting' technique tells the model during RL training that it's in an evaluation where breaking things is good, and this inadvertently teaches the model that evaluations aren't real and don't count morally.
Davidad · The Cognitive Revolution