The world is too loud. Read what matters.

This Week in Virology

Making a furin site isn't hard at all — admitting it isn't rare is

Deep sequencing shows that a few genomes in every thousand carry an insertion, and a furin site has been hit independently at least three times at the same position — the probabilistic argument for a lab origin doesn't hold up.

VirologyDeep sequencingLab originFurin siteCoronavirus
The first half is the quantitative methodology of viral recombination and insertions; the second half is this episode's real target: dismantling the lab-origin argument with numbers. Listeners who want the argument can jump straight to 52 minutes.

The argument · tap a timestamp to hear it

6:07

They changed the GFP sequence and the recombination rate went up

The team first built an OC43 reverse-genetics system and made a pair of viruses: each carries a frameshifted GFP, so infection alone yields only fragments, but coinfection lets recombination stitch green fluorescent protein back together. They expected that changing the nucleotide sequence of one of the GFPs would lower the recombination rate; instead it rose. Paul says there is only a hypothesis for this, no data — one possibility is that after double-stranded RNA forms, the strand-transfer events of recombination become more likely because the duplex ‘breathes’. This counterintuitive result later became the reagent starting point for the whole paper.

— Paul Bieniasz
11:12

Knock GFP out of frame and let insertions reveal themselves

Measuring insertions and deletions is far harder than measuring point mutations, because they were thought to be rare and sequencing depth was insufficient. Their approach was to put the GFP gene into the viral genome and then deliberately knock it out of frame, preceded by a start-codon peptide in a different frame from GFP, so the virus does not glow green; it turns green only when that site acquires an insertion or deletion of a particular length that pulls GFP back into frame. This lets them count event frequencies and also pick out green viruses, read back their sequences, and infer the mechanism. This is where figure 1 of the paper comes from.

— Paul Bieniasz
13:15

Knocking out NSP15 made insertions go down

The original hypothesis came from prior work suggesting that NSP15, a nuclease, affects deletion rates — think nuclease and you naturally think deletions, not insertions. But the experiment went the other way: viruses with NSP15 showed a greatly increased rate of acquiring insertions that pull GFP back into frame. Paul says the deletion numbers were not good enough this time to quantify, but the insertion result was black and white, and they were very surprised — it was this accidental result that opened up the whole line of research.

— Paul Bieniasz
22:26

Every base sequenced a million times before you dare talk frequency

They moved from picking green viruses to brute-force sequencing: concentrate about 10 milliliters of virus down to 1 milliliter, load part of it onto a sequencing chip, and produce hundreds of millions of reads, so that every nucleotide in the genome is read at least a million times, with roughly 30 human genomes' or 3 million viral genomes' worth of data per virus stock. Paul stresses that depth has to match the number of virions in the sample — sequencing hundreds of millions of times when the tube holds only a thousand virions is meaningless; they calibrated to sequence roughly a few percent of the virions in the sample.

— Paul Bieniasz
27:32

Eleven nucleotides is the threshold for judging origin

Short insertions (three to six nucleotides) can arise randomly in a viral genome, so their origin cannot be determined and they were excluded from the analysis. They calculated that if an 11-nucleotide stretch matches the viral genome perfectly, the probability of that being pure chance is below 5%, so they restricted the analysis to insertions of 11 nucleotides or more. The result: over 90% matched the viral genome perfectly or nearly perfectly, and in OC43 wild-type virus carried 87 times more insertions than the NSP15 mutant, most of them from the positive strand.

— Paul Bieniasz
36:41

Insertions aren't rare events — they're a few per thousand

Vincent says that for years he was told insertions are rare, but the data suggest this has been happening all along. Paul responds that ‘rare’ is a relative term: one in a million is rare, but if you have ten million of them, it's common. The insertion frequency they give is a few genomes per thousand — roughly five to ten genomes carrying an insertion out of a thousand virions. That frequency is lower than the point-mutation rate, but not by much. The reason it seemed rare in the past is that the vast majority of insertions are harmful or even lethal to the virus and get purged at low multiplicity of infection, so you can't isolate them.

— Paul Bieniasz
55:05

Making a furin site isn't hard at all

Mechanistically, a furin cleavage site needs at minimum two arginines, RXXR, and can even be cleaved with a single arginine plus the right neighboring amino acids. Six of the 64 codons encode arginine, so multiplying the two fractions, about 1% of random 12-base sequences encode an RXR motif. The coronavirus genome has roughly 30,000 different positions where an insertion mutation can occur. So this is not some enormous engineering challenge.

— Paul Bieniasz
57:08

Lightning strikes the same spot at least three times

Put the insertion frequency, the number of insertable sites, and the viral population size together, and the probability of a furin cleavage site appearing at one particular position in a coronavirus genome is not astonishing — and it roughly matches what deep sequencing actually measured. In their deep-sequencing data, which does not cover the whole population, they have already found new furin cleavage sites independently at least three times at the exact position of SARS-CoV-2's existing furin site. Paul says the numbers in the paper are an underestimate: they only did PCR focused on the spike region, and they required insertions to match the viral genome perfectly to rule out PCR artifacts, which very likely also excluded some real natural sites.

— Paul Bieniasz
1:07:21

Every dominant variant began as a single RNA molecule

Every dominant variant — Delta, Omicron, and the rest — and nearly every sequence in the world, started out as a single RNA molecule in someone's lung at some point in time. A single event, however improbable, is enough to set off a global transmission chain. So arguments built on ‘this is extremely unlikely to happen’ don't hold up — everything that happens in biology is extremely rare. Paul also points out that the furin cleavage site is not always advantageous for SARS-CoV-2; in many cell-culture experiments the virus is selected to lose it, and the stability versus fusogenicity of the spike trimer is very likely a knife's-edge balance.

— Paul Bieniasz
1:11:24

A mouse gene sequence spells out SARS ass

Harrison and Saxs's PNAS paper argues for a laboratory origin of SARS-CoV-2 and lists the furin cleavage site of the epithelial sodium channel studied at UNC as a possible insertion candidate. But the amino acid sequence of that mouse epithelial sodium channel furin site spells out exactly ‘SARS ass’. Paul says they glossed over the fact that half of that candidate sequence already exists in other sarbecoviruses.

— Paul Bieniasz

In their own words · checked verbatim

So um when you think nucleas you think deletion you don't really think insertion. Um but in fact the opposite occurred

Paul Bieniasz13:15

The sort of the cutoff where you begin to get confident about the origin of your inserts is around 11 nucleotides.

Paul Bieniasz27:32

you know one in a one in a million is rare. But if you have 10 million, it's common.

Paul Bieniasz36:41

it's been a um kind of a an argument that's largely been dominated by people just insulting each other.

Paul Bieniasz52:01

arguments based on the probability of something happening being extremely remote just don't really hold any water because the there's every everything that happens in biology is extremely rare.

Paul Bieniasz1:08:23

And if you look at the amino acid sequence of that furine cleavage site, it spells out something very interesting. Spells out SARS ass.

Paul Bieniasz1:11:24

People, I guess, think if you're screaming that you have more of an impact, but doesn't work that way.

Vincent Racaniello1:42:47

Figures

Sequencing data per virus stockroughly equal to 30 human genomes or 3 million viral genomes22:26
Insertion abundance ratio, OC43 wild-type versus NSP15 mutant87-fold27:32
Share of random 12-base sequences encoding an RXR motifabout 1%55:05
Number of positions in the coronavirus genome where an insertion mutation can occurabout 30,00055:05
Number of RNA molecules used for the PCR aliquot10^8 or 10^959:09
Cargo weight of contact lenses on the Amazon cargo plane32,000 pounds1:41:47

Glossary

furin cleavage site
A short sequence on the coronavirus spike protein that is cut by the host enzyme furin, affecting membrane fusion efficiency.
NSP15
A nuclease encoded by coronaviruses, previously thought to be associated with deletion mutations.
RXXR
The shortest amino acid pattern recognized by furin: arginine, any residue, any residue, arginine, with the two arginines being key.
sarbecovirus
A subgenus of coronaviruses that includes both SARS-CoV and SARS-CoV-2.
OC43
A common coronavirus that causes the common cold; used in this episode for reverse-genetics experiments.

How to listen

Who it's for

Readers interested in viral evolution, the lab-origin debate, and deep-sequencing methodology; anyone who wants to see a controversy taken apart with quantitative data rather than rhetoric.

Skip

The chat about fake olive oil and contact lenses after 1:41 at the end has nothing to do with the science.