Given enough context, written language is highly predictable, and this same statistical predictability is what modern LLMs exploit when they predict the next token.
CJ traces LLM next-token prediction back to Claude Shannon's 1950 letter-guessing experiment with his wife, which showed language has deep statistical structure.
transcript
CJ: his finding was, given enough context, the next letter is often nearly certain. Now, this means that language has deep statistical structure, and over 70 years later, that's pretty much what an LLM is doing. It's predicting what is coming next.
explains mechanism · 1