AI Didn't Collapse Because a Phase Transition Happened, Not Because of More Parameters
Physicist Nigel Goldenfeld uses the renormalization group and phase transitions to explain the mechanism by which AI still hasn't collapsed at a trillion parameters, and applies the same framework to answer why the genetic code is unique.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
The square-root law was wrong: the exponent simply isn't 1/2
Magnetization as a function of temperature should have followed a simple square-root law, but experiments gave an exponent of 0.3265136. From the late 1940s to the mid-1960s, physicists could not explain why the exponent wasn't 1/2, and this became a major puzzle in condensed matter physics. Nigel says that at the time no known method could account for why the exponent was not one half. The tool that could finally handle the question was the renormalization group, which came later.
— Nigel GoldenfeldIrreversibility made high-energy physics hard but made condensed matter work
Strictly speaking the renormalization group is a semigroup: it studies a physical system by repeatedly coarse-graining the spins. The crucial feature is irreversibility. There is a unique path from the microscopic to the macroscopic, but running it backwards cannot uniquely determine the microscopic state. That irreversibility is what made high-energy physics hard and what made condensed matter physics succeed. It is also why chemists need not worry about corrections from the mass of the top quark: the microscopic physics collapses into constants such as the mass of the proton. Without this coarse-grainability, science itself could not be done.
— Nigel GoldenfeldDimension is the key variable that decides how matter behaves
The Ising model has no phase transition in one dimension, which is why Ising himself gave up on it for a time; in 1944 Onsager solved the two-dimensional model, an achievement that has been compared to general relativity. Dimension is the key variable determining how matter behaves, and it is now even studied as a continuous variable. Understanding the effect of dimension is what makes it possible to understand why entirely new behavior appears in AI only in very large parameter spaces.
— Nigel GoldenfeldAI works not because of more parameters but because the system entered another state
Modern AI fits data with a trillion parameters and, by the overfitting bounds of traditional statistics, should long since have fallen apart, yet in practice it still works. Nigel and his collaborators found that there is a phase transition during learning, accompanied by critical exponents and data collapse, isomorphic to the superconducting phase transition. He stresses that the key word is ‘different’, not ‘more’: it is not that more parameters allow more data to be fitted, but that the system has entered a different physical state.
— Nigel GoldenfeldIn finite neural networks the phase transition vanishes and degenerates into ridge regression
Phase transitions exist strictly only in the limit of infinite systems. In finite systems — a finite network of neurons, for instance — the phase transition disappears and the behavior degenerates into ridge regression. Experimentally, you have to get within 10^-12 degrees of the critical temperature to see finite-size effects. This in turn explains why AI has to reach a trillion parameters: the scale itself is the condition for approaching the phase-transition limit.
— Nigel GoldenfeldThe genetic code is so close to optimal that it looks designed
Nigel reduces the questions about the genetic code to three puzzles: life appeared too early; why is there only one standard genetic code; and the code is close to optimal in its tolerance of errors, as though intelligently designed. This deeply puzzled Francis Crick. A genome is a DNA sequence, whereas the genetic code is the lookup table mapping 64 triplets onto 20 amino acids, and the two must be discussed separately.
— Nigel GoldenfeldThe code evolved fast because genes were passed sideways back then
Nigel's paper resolves all three puzzles using horizontal gene transfer. Early life did not rely on vertical inheritance; genes could spread sideways between organisms with no ancestral relationship, propagating rapidly the way network effects do, so the code could evolve and reach optimality in a short time. As complexity rose, the network effect shut off and evolution shifted to the vertical mode. This explains why the code is unique, why it is optimal, and why life appeared sooner than expected.
— Nigel GoldenfeldLife is a physical process, not a chemical one
Nigel defines life as a physical rather than a chemical process: life uses gradients in chemical potential to drive itself, and its purpose is to help a planet reach equilibrium. The article he co-wrote with Carl Woese is titled exactly that — ‘Life is Physics’ — and he holds that life can be implemented on different chemical substrates, silicon-based life included. The proposition pulls planetary science, statistical physics and biology into a single framework.
— Nigel GoldenfeldIn their own words · checked verbatim
And you're way, way into that regime where you're just fitting noise and the whole thing shouldn't work. It obviously shouldn't, and yet it does.
Nigel Goldenfeld0:00
The problem is that there's no known way or there was no known way to account for the fact that the member is not a half.
Nigel Goldenfeld2:05
It's not just that, well, you have more things and so you can fit more things. There's actually, it's different is the important thing, not the more.
Nigel Goldenfeld29:57
Why is there only one genetic code? Yeah. Okay. We call it the canonical genetic code. And there's minor variations, mainly to do with stop codons, but it's basically the same genetic code.
Nigel Goldenfeld44:29
If that was the process operative at the dawn of life, it wouldn't have happened this way. But, in fact, what happened was horizontal gene transfer, namely that genes can be transferred between organisms that are not related.
Nigel Goldenfeld48:43
The purpose of life is to help planets come into equilibrium.
Nigel Goldenfeld58:32
Figures
| Critical exponent | 0.3265136 | 2:05 |
| Dimension of the Ising model | no phase transition in one dimension, a phase transition in two | 11:22 |
| Year Onsager solved the two-dimensional Ising model | 1944 | 11:22 |
| AI parameter scale | a trillion | 26:56 |
| Number of genetic code triplets | 64 | 40:26 |
| Number of amino acid types | 20 | 40:26 |
| When the last common ancestor appeared | 3.8 billion years ago | 42:29 |
| Age of the meteorite | 4.35 billion years | 43:29 |
Glossary
- renormalization group
- A method that progressively coarse-grains and averages over microscopic degrees of freedom to study how physical laws change across scales.
- critical exponent
- The power-law exponent describing how a physical quantity varies with the control parameter near a phase transition, such as 0.3265136.
- data collapse
- Making data from different scales dimensionless so it falls onto a single curve, showing that the system has scale invariance.
- horizontal gene transfer
- The sideways spread of genes between organisms that are not in a direct parent-offspring relationship, as opposed to vertical inheritance.
- ridge regression
- Linear regression with L2 regularization, often the substitute description for a finite system that has no phase transition.
How to listen
Founders, investors and engineers who want to understand why AI hasn't overfit at a trillion parameters, plus researchers who want a physics framework for the origin of life.
The archaea and codon detail from 41:26–55:18 can be fast-forwarded; go straight to the ‘life is physics’ conclusion at 58:32.