The world is too loud. Read what matters.

Latent Space

AI can now build a virus; the defense side is still stuck on sequence alignment

The models that generate pathogenic sequences and the models that tell whether a sequence is pathogenic are essentially the same models, so the teams doing design are best positioned to do defense — but the defense side is far behind, with a baseline still stuck at sequence matching.

BiosecurityGenome modelsAI for ScienceArms raceDNA synthesis

The video won't play here. Listen to the audio instead:

High information density; the middle and later sections on the four layers of biodefense, the DNA synthesis pipeline, and "you can't patch a human" are the most valuable.

The argument · tap a timestamp to hear it

0:00

The teams doing design are best positioned to do defense

The dual mission Eric set for Radical Numerics is not a PR gesture but a technical judgment: the models that can generate pathogenic sequences and the models that can tell whether a sequence is pathogenic are essentially the same models. So the teams doing design are naturally best positioned to do defense, rather than outsourcing defense to a different group. He acknowledges there is an arms-race dynamic here, and that the defense side has always been far behind; what the company wants to do is pull the defense side up to the same level.

— Eric Nguyen
8:12

AI built the first genome from scratch

Evo's work generating a CRISPR-Cas system landed in Science, and Eric later gave a TED Talk. More important is what scientists showed last year: using a generative DNA model to generate the first genome from scratch. Humans had previously only been able to copy and paste fragments of known function, and had never built from scratch at the genome level. What Evo generated was a functional bacteriophage, that is, a virus. Eric says this is a turning point for the scientific community and the company — it means being able to build complete organisms that do not exist in nature, and it also means the corresponding potential for harm.

— Eric Nguyen
12:16

Omni's breakthrough came from alignment, not pretraining

Evo showed the potential of pretraining, but on human genetics it still lost to specialized DNA models, and the community questioned why large models were needed at all. Omni's difference lies in midtraining and posttraining: presenting tasks in the question-and-answer form scientists actually want, for example, given a wild-type and a mutant sequence, help me find the causal variant. Eric says this kind of form does not emerge naturally from pretraining. The results surprised them — Omni was not just competitive across many tasks the way Evo was, it began pushing the frontier in every subfield, especially variant effect prediction.

— Eric Nguyen
23:50

Most real disease variants are in non-coding regions

Traditional bioinformatics tools and statistical methods mainly focus on coding regions, and coding regions make up only about 1.5% to 2% of the genome. Eric says many, perhaps most, diseases actually fall in non-coding regulatory regions, where variants are harder to judge as pathogenic or not. This is exactly where Omni is strong: it can distinguish pathogenic mutations in non-coding regions, especially long-range ones. This is exactly what geneticists who have long been unable to use traditional tools want.

— Eric Nguyen
45:44

The biology experiment itself is doing the filtering

The mechanism of the SOLEX experiment is worth remembering: generate a large number of random sequences, turn them into aptamer switches, and read them out at high throughput with NGS — the more binding, the more reads, the higher the fitness; then mutate the winning sequences again and iterate. This experiment explored on the order of 10 to the 11th power sequences in parallel. The key is that "the biology of the experiment is actually doing the filtering" — the wet lab does the filtering, and the model reproduces that process in silico. This paradigm can be generalized to any biological experiment of "progressive refinement," as long as you can capture the intermediate states.

— Eric Nguyen
58:15

Stanford scientists said generating DNA was a stupid idea

Eric spent six months at Stanford on the idea of Evo "generating DNA," and almost every scientist he talked to thought it was a stupid idea. The objections were specific: humans themselves do not understand the rules of DNA, so why expect AI to learn them; you cannot even tell whether the generation is correct, so how do you give the model feedback; DNA has too many repetitive sequences, too noisy a distribution, no real rules, all junk. He admits he was just stubborn, felt "it just feels right," but at the time could not say what the use case was. The first time they trained Evo, they had no idea whether it would work.

— Eric Nguyen
1:12:48

The four pillars of biodefense

They break biodefense into four layers: detection and surveillance (detecting pathogen sequences from environmental samples such as airport nasal swabs and wastewater systems), attribution (judging whether the source is natural, foreign, or engineered in some country's lab), countermeasures (making antivirals or antibacterials), and deterrence (at the government level). They focus on the first three. The core criticism is this: the existing biodefense community mainly relies on sequence matching — aligning sequences against databases of known pathogens. This does not stop two things: new things that are not in the database, and designs that are deliberately obfuscated, that do not match in sequence space but have the same function.

— Eric Nguyen
1:25:15

You can't patch a human

The host points out the key difference between biosecurity and cybersecurity: in cybersecurity, a sufficiently strong model can in principle find every vulnerability and patch it, and as long as people keep updating, they are safe; but the genome is fixed, "you can't patch a human," so the defense surface is harder. Conversely there are two buffers: the intensity of motivation to launch an attack may be much lower, and the threshold from printing arbitrary DNA to building a successful virus — especially one that does not first kill the designer — is actually quite high.

— Host

In their own words · checked verbatim

We felt it was important as a lab that a team that was both building the design capabilities is actually also best suited for building the defense capabilities because they're basically the same models.

Eric Nguyen0:00

to build something from scratch and ground up had not been done before at the genome level

Eric Nguyen8:12

it turns out many, if not most of the diseases, are in these non-coding regions

Eric Nguyen23:50

But the key is the biology of the experiment is actually doing the filtering, right?

Eric Nguyen45:44

And crazy enough, most, almost every scientist at Stanford I talked to thought it was a stupid idea.

Eric Nguyen58:15

And I remember when the first result came back and it kind of gave us chills over like, oh, maybe there's something going on

Eric Nguyen1:00:25

And if you don't have a tool, then you can't do anything.

Eric Nguyen1:22:11

We have fixed genomes, right? So you can't patch a human.

Figures

Maximum context length supported by HyenaDNAone million3:02
Share of the genome that is coding regionabout 1.5% to 2%23:50
Annual global deaths from bacterial infectionsabout 2 million41:42
Number of sequences explored in parallel in the SOLEX experimentabout 10 to the 11th power45:44
Length of the Evo bacteriophage genomeabout 6000 base pairs54:07
Length of the human genomeabout 3 billion base pairs57:11
Current baseline of biodefensesequence-based matching, alignment matching1:22:11
Position of the defense sidefar, far lagging1:24:13

Glossary

GLM (Genome Language Model)
A large language model trained on DNA sequences that can both read and write DNA.
HyenaDNA
A DNA language model that replaces attention with convolution and pushes context to the million scale.
Evo
Radical Numerics' generative genome model, which once generated a functional bacteriophage from scratch.
Omni
Evo's successor model, which uses midtraining and posttraining to push the frontier on variant effect prediction.
biodefense
A defense system made up of four layers: detection and surveillance, attribution, countermeasures, and deterrence.
sequence matching
The existing defense baseline of aligning sequences against databases of known pathogens, which cannot stop new designs.

How to listen

Who it's for

Founders and investors working on AI for science, biosecurity, or genome models, and researchers who care about how the AI frontier lands in human health.

Skip

The listener farewell at 1:30:31 is generic and can be skipped.