Reinventing Entropy | Compression is Intelligence Part 1
3Blue1Brown
Jun 7, 2026
Claude Shannon's information theory reveals a profound link between predictive modeling and data compression. Modern machine learning achieves intelligence by approximating the most efficient possible compression of language, transforming our understanding of what cross-entropy loss actually signifies in model training.
Key insight: Shannon estimated the entropy of English to be about one bit per character, meaning human language is so predictable that, given enough context, it could theoretically be compressed to a single yes-or-no question per character.