A new benchmarking method using donor-specific assemblies found that Element AVITI™, particularly with UltraQ™ chemistry, achieved the lowest sequencing error rates of any platform tested—with the biggest advantage in low-complexity genomic regions.
Next-generation sequencing has transformed genomics, but a fundamental tension sits at the heart of accuracy benchmarking: how do you objectively measure errors in a sequencing platform when the tools used to measure those errors may themselves be biased by another platform's limitations?
Most sequencing accuracy assessments rely on established reference genomes or curated truth sets to identify errors. But these resources are not neutral. Reference genomes like GRCh38 were constructed predominantly using short-read, synthesis-by-sequencing (SBS)-based technologies. As a result, variants and sequence contexts that are difficult for SBS chemistry to resolve are systematically underrepresented or inaccurately annotated in the reference itself. When a short-read platform is evaluated against such a reference, errors intrinsic to the reference are invisible; variants that deviate from SBS assumptions may be misclassified as sequencing errors even when they reflect real genomic sequence. The reference, in effect, grades on its own curve.
This problem is not merely theoretical. As new sequencing chemistries enter the market, each with distinct mechanisms, error profiles, and accuracy claims, the need for a truly technology-agnostic benchmarking standard has never been more pressing.
To address this limitation, researchers at the Northwest Genomics Center (NWGC) and collaborators across the NIH's SMaHT (Somatic Mosaicism across Human Tissues) network developed a novel benchmarking framework built around what they call Donor-Specific Assemblies, or DSAs. This work was recently published in a pre-print: McGee et al. (2026). Short-Read Sequencing Benchmarking with Donor-Specific Assemblies, bioRxiv. DOI: 10.64898/2026.06.24.734333.
A DSA is a complete, diploid reference genome assembled de novo for each individual donor, using exclusively long-read sequencing data (PacBio HiFi and Oxford Nanopore Technologies) combined with Hi-C chromatin conformation data for chromosomal phasing and structural variant resolution. While Hi-C data employs short-read sequencing to map chromatin contacts and assist with chromosomal phasing and structural organization, no short-read sequencing contributes to base-level sequence determination in a DSA.
The consequence is significant: because a DSA's nucleotide sequence is built entirely from long-read consensus, it carries no inherent bias toward any short-read chemistry. Every base reflects the true sequence of that individual's genome, including variants and low-complexity regions that short-read assemblies historically struggle to resolve accurately.
"A DSA distinguishes between errors in the reference and errors in the data, a distinction that prior benchmarking frameworks could not reliably make."
When a short-read sequencing run is aligned to its donor-matched DSA rather than GRCh38, mismatches that arise from the reference being wrong rather than the sequencer being wrong are eliminated from the error calculation. What remains is a cleaner signal of platform-intrinsic sequencing error. With multiple DSAs generated for well-characterized cell lines (HG002 and COLO829BL), the group benchmarked nine short-read sequencing platforms simultaneously under identical analytical conditions, yielding the most comprehensive unbiased cross-platform comparison published to date.
With DSAs as the benchmarking foundation, the study generated a series of findings with meaningful implications for the field.
When raw, unfiltered data is assessed, single-nucleotide mismatch rates (SNMRs) across platforms are more similar than commonly assumed. Both Element AVITI and Illumina NovaSeq platforms show broadly comparable error rates at this stage, a finding that underscores the value of DSA-based benchmarking in establishing an honest baseline free of reference bias.
Figure 1-3. DSA-aligned single-nucleotide mismatch rates (SNMR) and indel mismatch rates (IDMR) across eight short-read sequencing platforms prior to quality filtering, shown for two cell lines (HG002 and COLO829BL). Values represent mismatches per 1,000 mapped bases. Platforms are ordered left to right: Illumina NovaSeq 6000, Illumina NovaSeq X, Element AVITI (Cloudbreak™), Element AVITI (UltraQ), PacBio Onso, MGI T7, Ultima UG100, and Ultima ppmSeq. Roche SBX-D chemistry not included in prefiltered results.
Applying a standard BQ>=30 quality filter, removing approximately the same proportion of reads across all platforms tested, produced a striking divergence in outcomes. Element AVITI with standard Cloudbreak chemistry and with UltraQ, a combination of advanced sequencing chemistry, enzymatic removal of deaminated bases during library preparation, and dark cycling to eliminate end-repair errors from data generation, both showed significantly lower error rates than either Illumina NovaSeq 6000 or Illumina NovaSeq X following standard quality filtering (BQ≥30). The magnitude of improvement seen with Element data substantially exceeded that observed with Illumina, despite comparable proportions of reads being removed.
This divergence reflects an important distinction in the nature of errors between platforms. A fraction of reads in Element data contain elevated error rates attributable to closely overlapping polonies, a known occurrence with randomly arrayed flow cells where a small number of adjacent clusters can interfere with one another during imaging. These reads produce detectably low base quality scores and are efficiently and specifically removed by standard BQ>=30 filtering, leaving behind a population of reads with extremely high accuracy.
By contrast, the more modest improvement seen in Illumina data after the same filtering step suggests that errors are distributed more broadly across the read population rather than concentrated in a readily identifiable subset; this pattern is more consistent with a systemic baseline error rate than a correctable, focal source of noise.
It is worth noting that quality filtering at or above BQ>=30 is a standard practice in pipelines demanding high accuracy, including variant calling in clinical research, somatic mutation detection, and low-frequency allele identification. The ability of a platform to improve substantially under this routine filter is therefore a practically meaningful performance attribute, not merely a statistical artifact. Notably, Element AVITI with UltraQ achieved the lowest post-filter error rates of any platform evaluated in the study, for both single-nucleotide variants and indels.
Figure 4 and 5. DSA-aligned SNMR and IDMR across nine short-read sequencing platforms following application of a BQ>=30 quality filter. Red triangles indicate the percentage of base calls retained after filtering for each platform and sample. The dramatic reduction in error rates for Element AVITI platforms relative to Illumina, despite comparable read filtering proportions, illustrates the concentration of Element errors in a small, identifiable subset of reads.
McGee et al. represents the first independent comparison of Roche's Sequencing-by-Expansion (SBX) chemistry, run on the Axelios instrument, against established short-read platforms. The inclusion of SBX in this analysis is particularly notable given that single-plex SBX reads carry a flat base quality assignment below standard thresholds; as such, the authors analyzed duplex sequencing data, in which, forward and reverse reads of the same molecule are combined to generate a higher-confidence consensus call.
In duplex mode, Roche SBX-D achieved competitive single-nucleotide mismatch rates, demonstrating the chemistry can produce high substitution accuracy when reads are processed through the duplex consensus workflow. However, indel error rates for SBX-D remained approximately an order of magnitude higher than those observed for the best-performing platforms in the study, even after duplex consensus calling. This finding persisted across both cell lines evaluated and represents a meaningful performance gap in a biologically critical error class.
Accurate detection of insertions and deletions is essential for a broad range of applications, and the requirement becomes more stringent as the target variant allele frequency decreases. In tumor profiling, where somatic mutations may be present in only a small fraction of cells within a heterogeneous sample, high indel background rates from the sequencing platform itself can obscure or mimic true low-frequency variants. High-value indels including EGFR exon 19 and 20 deletions, BRCA1 and BRCA2 frameshift mutations, and microsatellite instability markers are all routinely assessed in oncology research where sensitivity at low variant allele frequencies is a core requirement. Elevated intrinsic indel error rates, as observed with Roche SBX-D, present a practical barrier to the reliable detection of these variants in applications where they matter most.
When analysis was restricted to low-complexity regions (LCRs), genomic segments enriched for tandem repeats and homopolymers as defined by GA4GH benchmarking standards, performance differences between platforms widened substantially. Platforms employing flow-based and expansion-based chemistries showed the highest error concentrations in LCRs, consistent with known challenges in resolving homopolymer lengths and repeat structures using these approaches.
In low-complexity genomic regions, Element AVITI produced the lowest single-nucleotide and indel error rates of any of the nine sequencing platforms evaluated in the study. This performance is consistent with the known strengths of Avidite Base Chemistry™ (ABC™), which underpins the Element platform. The combination of the high specificity of avidite-based binding with the use of rolling circle amplification for polony generation makes ABC particularly well-suited to maintaining accuracy in the repetitive sequence contexts where other chemistries accumulate the most error. Both UltraQ and standard Cloudbreak chemistry significantly outperformed Illumina NovaSeq in LCRs for both variant types.
LCRs are disproportionately represented among actionable loci and are a well-documented source of false positives and false negatives in variant calling pipelines. For sequencing applications where LCR accuracy is non-negotiable, including research into repeat expansion disorder characterization, homopolymer-associated frameshift detection, and microsatellite instability assessment, these data provide a compelling basis for platform selection.
Figure 6. Comparison of SNMR and IDMR within low-complexity regions (LCR) versus non-LCR regions across all platforms, using DSA-aligned data with BQ>=30 filtering applied. LCRs are defined per GA4GH standards and include tandem repeat and homopolymer sequences lifted from GRCh38 to each donor-specific assembly. Element AVITI platforms demonstrate the lowest LCR error rates of any short-read platform evaluated.
Sequencing accuracy is not an abstract metric. As platforms improve, the sensitivity floor for variant detection drops; with it, the range of biological questions that can be meaningfully addressed expands. Applications that were once intractable are becoming routine across the field: comprehensive profiling of somatic mosaicism in healthy tissue, monitoring of circulating tumor DNA at ultra-low variant allele frequencies, and research into minimal residual disease (MRD) after cancer therapy. Each of these applications demands not just high mean accuracy, but consistent accuracy across difficult sequence contexts, including the low-complexity regions where many clinically significant variants reside.
The DSA framework introduced by McGee et al. provides the field with a more honest measuring stick, one that can keep pace with the rapid diversification of sequencing technologies and the increasingly demanding applications they are asked to support. As the sequencing landscape continues to evolve, studies like this one will be essential in helping researchers, and developers understand not just what platforms claim, but what they can reliably deliver.
References
For Research Use Only. Not for use in diagnostic procedures.