OurWord.
1 read 中文

The world is too loud. Read what matters.

Andrej Karpathy

Rarely updated, but every episode rewards frame-by-frame watching

Deep reads here1 / 15 episodes
Cadence~1 every 8.1 days
Latest2026-08-18
TopicsAI & Tech
PriorityT1
Andrej Karpathy 0716

Why cross-entropy loss won: it is training a text compressor

Cross-entropy loss was selected because it can be proved mathematically to be the only functional form that guarantees the loss is minimized exactly when the model's output equals the true distribution of the data — and that turns out to be equivalent to training the model into a text compressor approaching the Shannon limit.

8 Points 8 Quotes Information theoryCross-entropy