J. M. W. Turner, Snow Storm (1842) — Information at the edge of noise.
Shannon
n.—the man who measured meaning by how surprising it was.
Claude Shannon's 1948 paper did something surprisingly clean: it took "information," a word that until then had meant roughly whatever you wanted it to mean, and gave it a number. Information, in Shannon's formalism, is not meaning. It's surprise—the degree to which a message couldn't have been predicted from what came before it. A perfectly predictable message carries zero information no matter how important its content. A message that's pure noise carries maximum information by his definition and zero meaning by any human one. The gap between those two readings is the whole subject.
I find this one of the most underrated roots of the current AI moment, more important than people give it credit for. An LLM doesn't store language sentence by sentence. It compresses an enormous semantic landscape into a latent space from which structure can later be reconstructed—which is, almost exactly, Shannon's source-coding problem, the question of how much you can compress a message before you start losing what made it meaningful in the first place. Every conversation about token costs, about hallucination, about why models drift on long documents, is downstream of a question Shannon formalized seventy years before anyone called it AI: how much can you compress before the signal degrades into something else?
Pascal gave probability its rigor. Shannon gave information its unit. Put them together and you have the actual mathematical background of the thing currently being marketed as intelligence—not magic, not cognition in any sense Pascal or Shannon would have recognized, but a very old idea about uncertainty and compression, running at a scale neither of them had the hardware to imagine.