Blog - Next-Generation Sequencing | Element Biosciences

When "More Diversity" Is Just More Noise: What No-Diversity Amplicons Reveal About Sequencing Error

Written by Element Biosciences | Sep 11, 2026, 6:24:09 PM

A new preprint finds that standard sequencing pipelines can mistake error for real biological diversity, and the problem gets worse the deeper you sequence. AVITI™ produced 2.5-5x fewer of these false variants than Illumina NovaSeq, suggesting accuracy, not just depth, is what determines reliable diversity estimates. 

The problem: your diversity estimate might be an artifact

Amplicon-based metagenomics gives researchers the ability to catalog thousands of taxa from a single sample, transforming how we study microbial communities. Microbiomes, though, are just one example of a broader class of sequencing challenges: complex populations of related but distinct sequences, where accurately measuring diversity is central to understanding the underlying biology. Immune repertoires, viral quasispecies, and tumor cell populations all share this same basic structure.

In each case, the promise is the same: sequence deeply enough, and you can count every distinct molecule in the population.

But there's a catch that's easy to overlook. Every sequencing read carries some probability of error, and at the read depths modern instruments routinely generate, tens or hundreds of millions of reads per sample, even a very low per-base error rate can generate an enormous number of spurious, novel-looking sequences. A single true molecule can spawn dozens of "phantom" variants that differ by only one or two erroneous bases. Variant-calling algorithms are supposed to filter these out by modeling expected error rates, but if the model doesn't keep pace with real-world error accumulation, diversity estimates start reflecting sequencer noise as much as biology. That's a serious problem for any application (microbiome ecology, immune repertoire profiling, or rare-variant detection), where the whole point is to count how many distinct things are really there.

The approach: measuring "diversity" where none should exist

A new preprint, “Deep sequencing artificially inflates estimates of microbial diversity” from Henry, Laderman, and Bergelson at NYU and Simons Foundation tackles this problem head-on by asking a simple but clever question: what happens when you sequence something that shouldn't have any diversity at all?

The team built their analysis around "no-diversity amplicons": DNA regions amplified either from a single-copy host gene with minimal natural variation (GIGANTEA, in wild Arabidopsis thaliana) or from a synthetic spike-in sequence with a fixed, known identity. Because these amplicons are co-amplified and co-sequenced alongside real microbial targets (16S rRNA, ITS, gyrB, rpoB), any sequence variants ("ASVs") detected within them can only be technical artifacts; there is no legitimate biological source of diversity to explain them. In effect, the no-diversity amplicon acts as an internal control, a canary that reveals exactly how much of the "diversity" downstream pipelines report is actually manufactured by sequencing and bioinformatic error, rather than real biology.

The results: diversity inflation is real, exponential, and platform-dependent

The findings are striking. Using default DADA2 parameters in QIIME2, the authors found that ASV richness in these no-diversity controls climbed as high as several hundred distinct "variants" per sample, despite the true answer being one or two. This inflation wasn't a fixed background noise level; it increased exponentially once read depth exceeded roughly 10,000 reads per sample, and the same exponential pattern showed up in every real microbial amplicon tested (16S, ITS, gyrB, rpoB), with the magnitude of inflation varying by amplicon type. Notably, the effect wasn't limited to overall richness; the most abundant taxa within a sample showed disproportionately higher within-taxon diversity inflation as read depth increased, meaning the artifact could distort not just community-level counts, but conclusions about which specific organisms are "diversifying."

Truncating reads to shorter lengths helped reduce, though never fully eliminated, this inflation, since errors tended to accumulate toward the ends of reads. Neither rarefaction nor low-abundance filtering (two of the field's most common corrective techniques) meaningfully solved the problem on their own.

Head-to-head comparison: AVITI outperform NovaSeq on diversity accuracy

The most actionable result, though, came from a head-to-head platform comparison. When the same samples were sequenced on both an Illumina NovaSeq system and the AVITI Element platform (which uses Avidite Base Chemistry™ rather than standard sequencing-by-synthesis), the AVITI data showed substantially better results: no-diversity amplicon inflation was 2.5- to 5-fold lower than on NovaSeq at the highest read depths, and AVITI removed roughly twice as many erroneous sequences during quality filtering. Because AVITI's underlying error rate is lower and more evenly distributed across the read, its raw data required less aggressive truncation and filtering to reach an accurate diversity estimate, particularly for high-diversity soil microbiomes, where the correlation between truncation lengths (a proxy for robustness) was consistently stronger on AVITI than on Illumina.

Figure A: ASV richness of no-diversity amplicons across read depth for AVITI and NovaSeq (200 bp and 250 bp truncation lengths). Richness remains low and stable for both platforms at lower depths but increases sharply for NovaSeq above ~10⁵ reads/sample, while AVITI richness stays comparatively flat.

Why this matters beyond microbiome ecology

The authors' takeaway for microbiome researchers is direct:

  1. Avoid richness as a headline diversity metric.

  2. Favor shorter amplicons where possible.

  3. Where feasible, favor sequencing platforms, like AVITI, with lower and more consistent error profiles.

But the implications of this study extend well past counting microbial taxa.

Any sequencing application that depends on distinguishing a real, rare sequence from a sea of near-identical background reads runs into exactly the same fundamental limitation demonstrated here. Immune repertoire sequencing needs to resolve true receptor diversity from PCR and sequencing noise that can masquerade as novel clonotypes. Research to support early cancer detection and minimal residual disease (MRD) monitoring depend on identifying a handful of tumor-derived variant molecules diluted among millions of wild-type sequences, a signal that can be trivially swamped by sequencer error if base-calling accuracy isn't high enough.

In all of these settings, "sequence deeper" is not a substitute for "sequence more accurately"; past a certain depth, additional reads mostly add more opportunities for error, not more true signal.

This is where sequencing chemistry itself becomes a rate-limiting factor for scientific and clinical research progress. Platforms built around fundamentally higher per-base accuracy, like AVITI's Avidite Base Chemistry, don't just produce marginally cleaner data; they change what's achievable at all.

Detecting a single tumor variant at one-in-a-million frequency, resolving the true breadth of an immune repertoire, or reporting a microbiome diversity number that reflects biology rather than noise all hinge on the same underlying requirement: an error rate low enough that rare, real signal doesn't get buried under artifactual variation. As research for applications like MRD monitoring and early disease detection push toward ever-lower variant allele frequencies, technologies capable of that level of accuracy are likely to become not just a nice-to-have, but a prerequisite for the next generation of sequencing-based discovery.

 

references 

Deep sequencing artificially inflates estimates of microbial diversity, Lucas P. Henry, Eric Laderman, Joy Bergelson. bioRxiv 2026.09.01.748665; doi: https://doi.org/10.64898/2026.09.01.748665