Language has deep statistical structure, and an LLM is fundamentally doing the same thing Claude Shannon measured in 1950 — predicting what comes next given enough context.
CJ traces the core idea of LLMs back to Claude Shannon's 1950 letter-prediction game, establishing that LLMs are statistical models of language that predict what comes next.
transcript
CJ: Claude Shannon, the father of information theory, sat down with his wife Betty to play a game. And he showed her a passage of text with the next letter hidden, and she had to guess what came next, letter by letter. And he measured how often she was right. But what he was really measuring is how predictable is written language. And his finding was, given enough context, the next letter is often nearly certain.
provides context · 1