AI for Science Is Stuck on the Last 0.001 Points, Not on Compute
ERA can churn out thousands of candidate notebooks, but it can't close the gap in the last 30 places — because the inner loop is a machine, and the outer loop still needs a human to judge ‘you got this paper wrong’.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
ERA takes chat as input and returns a scored notebook
ERA's product form isn't having people write config files — you just start chatting. Because mapping a scientific problem onto a scorable task is itself non-obvious, Michael Brenner wrote a dedicated agent to conversationally help you define which score to optimize. Once the conversation ends, it generates a Python notebook underneath containing a function that produces a score, and then the system starts mutating that notebook in a ‘clever way’, continually proposing new code to push the score higher.
— John PlattIt doesn't always pick the best candidate
ERA runs Monte Carlo Tree Search internally, keeping hundreds or thousands of notebook candidates at once in a tree-shaped pool. Selection uses the upper confidence bound from reinforcement learning: estimate the 95th percentile outcome of a mutation and pick the one with the highest optimistic upper bound. So it doesn't always greedily pick the currently best-performing notebook; sometimes it picks the fifth best, in order to explore more efficiently. It has also tried combining the ideas of two candidates to generate a third.
— John PlattEvolutionary coding was stuck for decades because random mutation is useless
Platt says people have been doing evolutionary coding since the 1970s — mutating Lisp code and the like — but it never took off, because random mutation in code space is basically worthless: just like DNA, the vast majority of mutations are harmful. ERA works because the inner loop is itself an AI that knows which direction to try, which amounts to having a meaningful gradient underneath. The price is that it also overfits, and it also goes off the rails in ‘genie wish’ fashion.
— John PlattPredictive models and descriptive models are two different things
John Platt splits scientific modeling into two kinds: a predictive model only seeks the lowest error rate on some dataset; a descriptive model is what science actually wants — it contains some internal description of reality, so it can extrapolate. His example is Newton: an apple falling can be fitted, but gravity isn't about apples, and if you take a model that only fits apples and ask it about planets, it has no planetary data and knows nothing. He admits the line is blurry, because the scientist themself also constrains the model with a pile of known facts and intuitions.
— John PlattThis is a power tool, and it will cut your fingers off
Faced with multiple hypothesis testing and overfitting, Platt's answer isn't a new method but a harsher old one: because ERA is a power tool, you have to be stricter than before, with a holdout buried so deep you never look at it yourself. He states plainly that no system yet discovers entirely new physics or entirely new science, and scientists are still responsible for judging whether a conclusion is descriptive.
— John PlattAI is stuck on the last 0.001 points
Ira pushed the score very close, but it didn't close the gap in roughly the last 30 places, because nobody goes and scrapes off that last 0.001 points. This exposes the system's real structure: the inner layer is a machine running for hours to produce candidates, and the outer layer is a human making judgments like ‘you got this paper wrong’. Platt compares it to a graduate student who never sleeps and is extremely eager — you tell it a direction, it goes and tries, but when it gets stuck a human has to change the approach.
— John PlattThe counterfactual model stalled the team for two years
To judge which contrail to avoid, you need to know how much warming it actually causes, and that's a counterfactual problem: you can't enter the universe where the contrail doesn't exist. The team produced a counterfactual model that worked for outgoing longwave radiation, but they could never build one for the part that reflects sunlight, and it stalled them for two years. They even self-tested with a dataset where contrails were artificially injected, and their own attempt couldn't pass their own test. In the end ERA found a simple model that searches over all confounders, passed the test, and cracked the problem.
— John PlattA wildfire is best put out before it's room-sized
Google has a crisis resilience effort, and one project in it is called Firesat: a wildfire is extremely easy to extinguish when it's only as big as ‘this room’, and once it reaches an acre it's far harder, while the growth from room-sized to an acre may be exponential. The plan is a low-Earth-orbit satellite constellation using mid-wave IR (which corresponds to blackbody radiation, exactly the temperature of fire) for detection; about 50 to 80 satellites could cover the globe, spot a fire about 5 meters across, and raise an alarm within 15 to 20 minutes. The partner is the nonprofit Earth Fire Alliance, the satellites are built by Muon Space, and one prototype has already launched.
— John PlattIn their own words · checked verbatim
But the reason why it just hasn't taken off is that random mutation in code space is pretty much worthless.
John Platt18:38
A descriptive model is actually what science is trying to get to, which is, okay, it should be able to extrapolate because it has sort of the physics or the actual, some description of reality that's captured within it.
John Platt22:44
Because it's a power tool, it can, I shouldn't probably say, it can slice your fingers off.
John Platt31:59
any metric that becomes a target is no longer good as a metric
John Platt35:02
So it's almost like having a hyper eager grad student or something who doesn't sleep. And you sort of tell it things and you and you sort of guide it around.
John Platt40:23
But would have happened if there hadn't been a contrail there. That's a very difficult thing to estimate because you can't access the universe where that didn't happen.
John Platt48:08
But when it's like climate and it's open and it's non-stationary or you have to make these big extrapolations, you have to be much more cautious.
John Platt1:06:54
it's very easy to put out a wildfire the size of this room. But even if it's like an acre, it gets much, much harder.
John Platt1:24:27
Figures
| ERA default number of parallel search branches | 10 | 12:27 |
| Quantile estimated by UCB | 95th percentile | 10:24 |
| Contrails' share of anthropogenic global warming | about 1% | 43:31 |
| Local contrail forcing over Europe | about 1 watt per square meter | 44:47 |
| Global mean anthropogenic warming forcing | about 3 watts per square meter | 44:47 |
| Protein predictions released by AlphaFold | about 6 billion | 1:06:54 |
| Low-orbit satellites needed for Firesat | about 50 to 80 | 1:24:27 |
| Smallest fire Firesat can detect | about 5 meters across | 1:24:27 |
| Firesat alarm latency | within 15 to 20 minutes (requires about 80 satellites) | 1:24:27 |
Glossary
- ERA
- Google's AI for science system, which uses evolutionary search to automatically generate and optimize scientific notebooks.
- Monte Carlo Tree Search
- An algorithm that balances exploration and exploitation in a huge search space, commonly used in games and code generation.
- upper confidence bound
- A reinforcement learning action-selection strategy that uses optimistic estimates to drive exploration.
- contrail
- Ice-crystal clouds left behind by aircraft flying in ice-supersaturated regions at high altitude, with a warming effect.
- NISQ
- Noisy Intermediate-Scale Quantum, the stage current quantum computing is at.
- broom sensor
- A sensor that sweeps across the spectrum in one direction as the satellite moves, rather than being a snapshot camera.
How to listen
Engineers and researchers working on AI for science, climate modeling or quantum computing, plus founders who want a look inside Google's internal toolchain.
If you only care about AI applications, the asteroid and Oscar anecdotes after 1:47 can be skipped.