AI can now build a virus; the defense side is still stuck on sequence alignment
The models that generate pathogenic sequences and the models that tell whether a sequence is pathogenic are essentially the same models, so the teams doing design are best positioned to do defense — but the defense side is far behind, with a baseline still stuck at sequence matching.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
The teams doing design are best positioned to do defense
The dual mission Eric set for Radical Numerics is not a PR gesture but a technical judgment: the models that can generate pathogenic sequences and the models that can tell whether a sequence is pathogenic are essentially the same models. So the teams doing design are naturally best positioned to do defense, rather than outsourcing defense to a different group. He acknowledges there is an arms-race dynamic here, and that the defense side has always been far behind; what the company wants to do is pull the defense side up to the same level.
— Eric NguyenAI built the first genome from scratch
Evo's work generating a CRISPR-Cas system landed in Science, and Eric later gave a TED Talk. More important is what scientists showed last year: using a generative DNA model to generate the first genome from scratch. Humans had previously only been able to copy and paste fragments of known function, and had never built from scratch at the genome level. What Evo generated was a functional bacteriophage, that is, a virus. Eric says this is a turning point for the scientific community and the company — it means being able to build complete organisms that do not exist in nature, and it also means the corresponding potential for harm.
— Eric NguyenOmni's breakthrough came from alignment, not pretraining
Evo showed the potential of pretraining, but on human genetics it still lost to specialized DNA models, and the community questioned why large models were needed at all. Omni's difference lies in midtraining and posttraining: presenting tasks in the question-and-answer form scientists actually want, for example, given a wild-type and a mutant sequence, help me find the causal variant. Eric says this kind of form does not emerge naturally from pretraining. The results surprised them — Omni was not just competitive across many tasks the way Evo was, it began pushing the frontier in every subfield, especially variant effect prediction.
— Eric NguyenMost real disease variants are in non-coding regions
Traditional bioinformatics tools and statistical methods mainly focus on coding regions, and coding regions make up only about 1.5% to 2% of the genome. Eric says many, perhaps most, diseases actually fall in non-coding regulatory regions, where variants are harder to judge as pathogenic or not. This is exactly where Omni is strong: it can distinguish pathogenic mutations in non-coding regions, especially long-range ones. This is exactly what geneticists who have long been unable to use traditional tools want.
— Eric NguyenThe biology experiment itself is doing the filtering
The mechanism of the SOLEX experiment is worth remembering: generate a large number of random sequences, turn them into aptamer switches, and read them out at high throughput with NGS — the more binding, the more reads, the higher the fitness; then mutate the winning sequences again and iterate. This experiment explored on the order of 10 to the 11th power sequences in parallel. The key is that "the biology of the experiment is actually doing the filtering" — the wet lab does the filtering, and the model reproduces that process in silico. This paradigm can be generalized to any biological experiment of "progressive refinement," as long as you can capture the intermediate states.
— Eric NguyenStanford scientists said generating DNA was a stupid idea
Eric spent six months at Stanford on the idea of Evo "generating DNA," and almost every scientist he talked to thought it was a stupid idea. The objections were specific: humans themselves do not understand the rules of DNA, so why expect AI to learn them; you cannot even tell whether the generation is correct, so how do you give the model feedback; DNA has too many repetitive sequences, too noisy a distribution, no real rules, all junk. He admits he was just stubborn, felt "it just feels right," but at the time could not say what the use case was. The first time they trained Evo, they had no idea whether it would work.
— Eric NguyenThe four pillars of biodefense
They break biodefense into four layers: detection and surveillance (detecting pathogen sequences from environmental samples such as airport nasal swabs and wastewater systems), attribution (judging whether the source is natural, foreign, or engineered in some country's lab), countermeasures (making antivirals or antibacterials), and deterrence (at the government level). They focus on the first three. The core criticism is this: the existing biodefense community mainly relies on sequence matching — aligning sequences against databases of known pathogens. This does not stop two things: new things that are not in the database, and designs that are deliberately obfuscated, that do not match in sequence space but have the same function.
— Eric NguyenYou can't patch a human
The host points out the key difference between biosecurity and cybersecurity: in cybersecurity, a sufficiently strong model can in principle find every vulnerability and patch it, and as long as people keep updating, they are safe; but the genome is fixed, "you can't patch a human," so the defense surface is harder. Conversely there are two buffers: the intensity of motivation to launch an attack may be much lower, and the threshold from printing arbitrary DNA to building a successful virus — especially one that does not first kill the designer — is actually quite high.
— HostIn their own words · checked verbatim
We felt it was important as a lab that a team that was both building the design capabilities is actually also best suited for building the defense capabilities because they're basically the same models.
Eric Nguyen0:00
to build something from scratch and ground up had not been done before at the genome level
Eric Nguyen8:12
it turns out many, if not most of the diseases, are in these non-coding regions
Eric Nguyen23:50
But the key is the biology of the experiment is actually doing the filtering, right?
Eric Nguyen45:44
And crazy enough, most, almost every scientist at Stanford I talked to thought it was a stupid idea.
Eric Nguyen58:15
And I remember when the first result came back and it kind of gave us chills over like, oh, maybe there's something going on
Eric Nguyen1:00:25
And if you don't have a tool, then you can't do anything.
Eric Nguyen1:22:11
We have fixed genomes, right? So you can't patch a human.
Host1:25:15
Figures
| Maximum context length supported by HyenaDNA | one million | 3:02 |
| Share of the genome that is coding region | about 1.5% to 2% | 23:50 |
| Annual global deaths from bacterial infections | about 2 million | 41:42 |
| Number of sequences explored in parallel in the SOLEX experiment | about 10 to the 11th power | 45:44 |
| Length of the Evo bacteriophage genome | about 6000 base pairs | 54:07 |
| Length of the human genome | about 3 billion base pairs | 57:11 |
| Current baseline of biodefense | sequence-based matching, alignment matching | 1:22:11 |
| Position of the defense side | far, far lagging | 1:24:13 |
Glossary
- GLM (Genome Language Model)
- A large language model trained on DNA sequences that can both read and write DNA.
- HyenaDNA
- A DNA language model that replaces attention with convolution and pushes context to the million scale.
- Evo
- Radical Numerics' generative genome model, which once generated a functional bacteriophage from scratch.
- Omni
- Evo's successor model, which uses midtraining and posttraining to push the frontier on variant effect prediction.
- biodefense
- A defense system made up of four layers: detection and surveillance, attribution, countermeasures, and deterrence.
- sequence matching
- The existing defense baseline of aligning sequences against databases of known pathogens, which cannot stop new designs.
How to listen
Founders and investors working on AI for science, biosecurity, or genome models, and researchers who care about how the AI frontier lands in human health.
The listener farewell at 1:30:31 is generic and can be skipped.