The world is too loud. Read what matters.

张小珺·商业访谈录

AI Is Just a Wright Flyer That Barely Left the Ground, and We Still Can't Control It

Deep learning founder Terry Sejnowski says what large models lack isn't language but a basal ganglia — no value function, no long-term memory, no emotion — and all of those brain regions can be added. The real risk isn't scale, it's that no one knows what it will grow into.

Deep learningLarge modelsCognitive scienceAI regulationBrain science
The first half is first-hand recollection of deep learning history (Boltzmann machines, Minsky, Chomsky); the second half is a mechanistic judgment about what large models are missing. High information density, but it takes a little patience.

The argument · tap a timestamp to hear it

5:10

The Boltzmann machine is an existence proof that Minsky was wrong

In their perceptron book, Minsky and Papert proved the limits of the single-layer perceptron and concluded by asserting that no one could extend the learning rule to multiple layers. Sejnowski and Hinton's approach was to change the units from deterministic 0/1 into probabilistic, fluctuating units, so the Boltzmann machine became an existence proof that "Minsky and Papert were wrong." Its key difference from backpropagation is that it is entirely local: no error needs to be sent back; you just compute the correlation between inputs and outputs under two conditions — input present and input withdrawn — and subtract them. They called the input-withdrawn phase the sleep phase. The cost is that you have to wait for the whole network to reach equilibrium and compute average correlations, which consumes far more compute than backpropagation.

— Terrence Sejnowski
8:10

The Boltzmann machine requires the whole network to be globally coherent

Sejnowski explains the Boltzmann machine's limits using the physics of phase transitions: the more layers, the more the input travels up and back down, and the whole network must become a single coordinated state, which he calls coherence — like how the entire system becomes coherent near the critical point where water turns to steam. This is both its elegance and the reason it is slow: that coherence has to be built up layer by layer. He also stresses that the Boltzmann machine was already a deep network with many hidden layers in the 1980s, it just wasn't called that at the time, and that it can do supervised learning as well as learn the probability distribution of the input, not merely learn a classification mapping.

— Terrence Sejnowski
17:20

Minsky's ping-pong robot had no budget for vision

Sejnowski tells a story he verified face-to-face with Minsky: the AI lab's first DARPA grant was to build a robot that could play ping-pong, and only after getting the money did they realize no funding had been requested to write the vision program, so it was handed to a graduate student as a summer project. At the 50th anniversary celebration he asked Minsky whether this was true, and Minsky corrected him: not a graduate student, an undergraduate. From this Sejnowski draws the fundamental problem with symbolism: the bigger the problem, the longer the program, and writing programs is extremely expensive and doesn't scale — even billions of dollars for billions of lines of code wouldn't solve it.

— Terrence Sejnowski
23:24

The Wright brothers studied gliding, not flapping

Sejnowski uses the Wright brothers to rebut "you can't learn to build a plane by watching birds": what they observed was not flapping but gliding, when birds don't flap, and they even built their own wind tunnel to design wings. They learned the principle, not the details — feathers are a material that is light with a large surface area, so they built a frame from wood and covered it with canvas, achieving the same high surface area and low weight. At Kitty Hawk they only got ten or twenty feet off the ground and flew about half a mile, but that was a proof of principle. He argues the hardest problem was actually control, later solved with wing flaps, and that when a bird turns it is twisting its wings.

— Terrence Sejnowski
28:27

The finite speed of light constrains both supercomputers and brains

Sejnowski visited a supercomputing center in Texas, where they told him the hardest part is that the speed of light isn't infinite — a nanosecond travels about a foot, and a nanosecond is exactly the clock cycle at gigahertz, so the latency of the wires between cores becomes critical. He points out that nature faces the same problem: there is conduction delay between neurons, and nature has already solved it. This is exactly the work he is doing now — figuring out how nature solves latency, then applying it to the design of massively parallel architectures, and continuing to scale up.

— Terrence Sejnowski
33:40

He asked Minsky in public: are you the devil

At the dinner of the 50th anniversary celebration, Sejnowski raised his hand and asked Minsky: some in the neural network community think you are the devil, because you stalled progress for decades — are you the devil. He says what made him angry was not Minsky's views but his attitude toward students — Minsky finally stood up and told the students shame on you, saying that doing applications was failure and that they should be doing artificial general intelligence. Sejnowski considers this abusing students and harming them. When asked, Minsky launched into a long talk about complexity and the nature of computation, and Sejnowski cut him off saying he had asked a yes-or-no question, and Minsky finally said yes, I'm the devil.

— Terrence Sejnowski
37:57

Self-supervision makes training data semi-infinite

Sejnowski argues the real breakthrough of generative models is self-supervision: previously, object recognition required manually labeling images, which was costly and limited, and the bigger the network the more data it needed, so data volume in turn constrained network size. Self-supervision labels nothing; you just train it to predict the next word in a sentence, and you can feed it every sentence from everywhere, so training data becomes semi-infinite and the scale limit disappears. He acknowledges there are still many important technical details, but this is the one that surprised him — no one predicted large models would have these capabilities, and what astonished him most is that their English grammar is so perfect it doesn't seem human.

— Terrence Sejnowski
45:08

What large models lack is a basal ganglia, not language

Sejnowski says he is writing a book about the new language models, and in it there is a point he thinks others have missed: you can't blame GPT-3, it has no parents. The part of the human brain responsible for reinforcement learning is the basal ganglia beneath the cortex, which learns sequences of actions that achieve goals and needs the world to feed back what is good and what is bad; AlphaGo is precisely two parts — a deep learning network plus a reinforcement learning engine that assigns a value to every position. Large models have no value function, and that is the missing link. He also says emotion would be easier to install than language, and so would long-term memory — you just simulate the hippocampus, and the brain has hundreds of such components, whereas today's large models have only the cortical part.

— Terrence Sejnowski

In their own words · checked verbatim

it proved that Minsky and Papert were wrong

Terrence Sejnowski5:10

as the problem gets bigger and bigger if you try to solve it with writing a computer program the program gets bigger and bigger and that's very labor intensive

Terrence Sejnowski18:20

i thought i was angry not because of what he said but the way he said it to his students

Terrence Sejnowski33:40

the only thing we can be absolutely sure of is that it's not human it's not human it's an alien

Terrence Sejnowski43:01

You can't blame GB3. They didn't have parents.

Terrence Sejnowski45:08

what you should be capping is a capability not the size

Terrence Sejnowski53:23

the most important things that will come out of it are ones that we can't even imagine that we don't even know about yet

Terrence Sejnowski1:00:38

if anybody tells you that oh it can't do general artificial intelligence right well just wait for tomorrow it's a moving target

Terrence Sejnowski1:01:43

Figures

Order of magnitude of connections/parameters in the human brain10 to the 14th to 15th power, about a thousand times more than current models12:12
Order of magnitude of current model parametersup to about a trillion11:12
Error rate reduction of AlexNet on ImageNet20%13:16
Processing layers in primate visual cortexabout 12 layers13:16
Scale of the NeurIPS conferenceabout 16,000 in person, about 3,000 online, roughly 19,000 total50:12
Cost of training a large modeltens of millions of dollars, months of machine time56:28
Latency corresponding to the speed of lighta nanosecond travels about a foot28:27

Glossary

Boltzmann machine
A multi-layer network whose units are made probabilistic and fluctuating, learning via local correlations; a precursor to deep networks.
self-supervision
Constructing training targets from the data itself without manual labels, such as predicting the next word in a sentence.
basal ganglia
The brain region beneath the cortex responsible for reinforcement learning; it learns action sequences that achieve goals and needs external feedback.
hippocampus
The brain region responsible for forming long-term memories; large models currently have no corresponding component.
stochastic parrots
A derogatory label critics give large models, meaning they merely regurgitate statistical patterns without understanding meaning.

How to listen

Who it's for

Engineers building models and AI products, investors tracking long-term AI judgments, and anyone who wants first-hand early deep learning history.

Skip

The opening section about his personal education and physics background can be fast-forwarded.