SNP arrays vs. genotyping-by-sequencing: A technical comparison for livestock breeding programs

    August 3, 2026
    Justin Kos, PhD

    Sign up for our newsletter

    Join our scientific community to stay up to date with Element news, insights, and product updates.

    This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

    Agrigenomics | Educational Series

    Genotyping method selection is not a settled question in livestock breeding. Three approaches—SNP microarrays, targeted sequencing (hybridization capture), and low-pass whole-genome sequencing with imputation—are all in active commercial and research use today, each with distinct chemistry, cost structure, and error profile. This piece lays out how each method works, where it holds up, where it doesn’t, and a less-discussed dimension of the comparison: what happens at the small fraction of sites where a microarray and a sequencing-based method disagree.

    What genotyping is actually measuring

    Genotyping infers which allele, or alleles, an animal carries at a defined set of genomic positions. Those calls feed parentage verification, marker-assisted selection at major-effect loci, and genomic evaluations (EBVs/GEBVs) that combine genotype with pedigree and phenotype data. The accuracy and density of the underlying genotype call bounds the accuracy of everything downstream including imputation, evaluation reliability, and ultimately genetic gain.

    How a lab arrives at that call differs substantially by method.

    Three methods, three cost-benefit profiles

    1. SNP microarrays

    Microarrays (Illumina Infinium, Affymetrix Axiom, ThermoFisher GeneTitan, and similar chemistries) query a fixed panel of SNP sites via allele-specific probe hybridization and signal detection. The genotype call is an inference from fluorescence intensity clustering rather than a direct read of the base.

    Strengths: mature and high-throughput, low per-sample cost at scale, deep integration with existing LIMS and national database submission formats (CDCB, ICAR), and large historical reference populations for calling and imputation.

    Limitations:

    • Fixed content: Adding or changing markers means redesigning and re-manufacturing the array — typically a 12 to 18-month cycle.

    • Ascertainment bias: Panel content is chosen from a discovery population, so minor allele frequency and marker informativeness both degrade for breeds or crossbreds underrepresented in that discovery set.

    • Indirect calling: Because the assay infers genotype from signal intensity rather than sequence, calls can be destabilized by structural variation, indels, or SNPs adjacent to the probe binding site, conditions that array quality checks can catch when they produce an ambiguous reading, but not when they produce a clean reading at the wrong genotype.

    • Manifest/coordinate errors: Probe-to-genome coordinate assignments can be off. Without a sequence-level readout, this class of error is generally undetectable without an orthogonal method.

    2. Targeted sequencing/hybridization capture

    Rather than an intensity readout, targeted sequencing methods (hybrid capture panels run on NGS) isolate regions of interest via bait-based capture, then sequence those fragments directly, base by base.

    Strengths: direct base-level evidence at each site, with read depth, base quality, and phasing information attached; the ability to detect variants adjacent to or overlapping a primary target that would otherwise obscure an array signal; and panel content that updates computationally (new bait design) rather than requiring new hardware, typically weeks rather than a year-plus.

    Limitations: a comparatively newer workflow for agrigenomics at this scale, requiring a provider with capture and NGS expertise; and calling accuracy depends on the maturity of the secondary analysis pipeline including variant calling, QC thresholds, and coverage uniformity across the panel.

    3. Low-pass whole-genome sequencing (lpWGS) with imputation

    This approach sequences the whole genome at low coverage (often well under 1x per sample) and statistically imputes a dense genotype using a haplotype reference panel.

    Strengths: whole-genome scope, including regions no fixed panel targets; cost that scales favorably with volume; accuracy that improves as reference panels grow; the ability to discover new markers; and the ability to reanalyze banked data with new markers as phenotype/genotype knowledge evolves.

    Limitations: accuracy is a direct function of reference panel depth and breed representation. Well-characterized breeds (e.g., Holstein) impute well; rarer breeds or novel crosses impute with more uncertainty, because there is less haplotype context available to triangulate from.

    No universal answer but the landscape is shifting

    There is no single best method across all breeds, budgets, and use cases; arrays remain a popular choice for well-characterized breeds with stable content needs. Several trends are changing the calculus, though: the accelerating pace at which new trait markers need to be added (feed efficiency, methane emissions, disease resistance), increasingly competitive sequencing costs, and formal acceptance of sequencing-based methods by evaluation bodies

    In June 2025, the Council on Dairy Cattle Breeding (CDCB) certified its first sequencing-based laboratory for submission to the U.S. national evaluation system, using a hybrid low-pass/targeted-sequencing approach, the first time genotyping-by-sequencing has been used in a national dairy genetic evaluation system anywhere in the world.1 Sequencing-based genotyping is shifting from a research-only tool and gaining acceptance as an evaluation-ready option.

    The accuracy question arrays don’t surface: a deep dive

    Cost and nimbleness differences between arrays and sequencing are well documented. Less discussed is what a same-sample, orthogonal-method comparison shows about array accuracy specifically.

    In a comparison of same-sample bovine genotyping data: Illumina array calls versus a Twist Bioscience 100k bovine capture panel, sequenced on an Element AVITI™ platform and analyzed with Curio, the two methods agreed at 99.5% of sites. That concordance rate is consistent with published array-to-array and array-to-sequencing comparisons.2,3 What’s revealing is what’s happening at the remaining 0.5%.

    SeqvsArraysFig1

    Figure 1. Alignment of NGS and array data along a 4 Mb region of the Bos taurus genome shows 5 out of 31 heterozygous calls have a concordance issue.

    Two discordance patterns recur

    Adjacent-SNP signal interference: At one discordant site, sequencing showed a heterozygous call fully supported by 29 reads, with three additional SNPs adjacent to the target position. The array called the site homozygous.

    Extra variation near the probe is a plausible explanation, since it can weaken the signal from one allele though confirming that mechanism at any individual site requires an independent method. What makes this class of discordance hard to catch comes down to how an array decides on a call. The instrument measures brightness at each site and compares it to the brightness patterns it expects for each genotype. Quality checks then ask a single question: does this reading look like one of the expected patterns, or does it fall somewhere in between? Readings that land in between get flagged. But a reading that lands squarely on the wrong pattern passes every check, because there is nothing unusual about it. When nearby variation suppresses one allele entirely, a heterozygote produces a brightness reading that looks exactly like a normal homozygote, and is called with full confidence. Seqvsarrays_figure2

    Figure 2. NGS data correctly calls the highlighted position as heterozygous, with 29 supporting reads and 3 adjacent supporting SNPs.

    Manifest/coordinate error: At a second site, the sequence-confirmed variant position was one base off from the array’s assigned coordinate. Whether this originated as a manifest error or as signal variability from a nearby SNP, it is a class of error that generally requires an independent, position-anchored readout to catch. Array workflows do not, on their own, verify their coordinate assignments against the underlying sequence.

    It is worth noting that no-calls and mis-calls sit on opposite sides of the same confidence threshold. When nearby variation partially disrupts a probe, the reading falls between the expected patterns and the array declines to call it. This is how the array is expected to function. However, when there is complete disruption, the reading lands cleanly on the wrong pattern and is called with confidence.

    Seqvsarrays_figure3

    Figure 3. Coordinate discrepancies of this kind require an independent sequence-level readout to identify. Adjacent SNPs may also contribute to inconsistent array signal intensity.

    Coming back to the broader 4 Mb region, most of the missed sites were heterozygous, the genotype class most vulnerable to intensity-based ambiguity.

    None of this means arrays are broadly unreliable; 99.5% concordance is a high bar, and most sites are unaffected. But the discordant sites are not randomly distributed, they cluster in SNP-dense regions, which is often exactly where allelic complexity matters for trait mapping or biomarker linking.

    One example of an integrated sequencing-based workflow

    Because targeted sequencing produces base-level, orthogonally verifiable data, it can directly address the failure mode described above and it can confirm a call independently of the panel’s expected signal pattern. Several vendors and integrators offer capture-based genotyping workflows; the one behind the comparison above is a joint Twist Bioscience/Element Biosciences™/Curio pipeline, run sample to report:

    • Library preparation: Twist FlexPrep, an auto-normalized, early-pooling library prep built for high multiplexing.

    • Capture: Twist hybridization probes, compatible with solution-phase capture or Trinity™ on-flow cell enrichment.

    • Sequencing: AVITI or AVITI24™, at 1x300 or 2x150, using Cloudbreak Freestyle™ or Trinity chemistry.

    • Analysis: direct data transfer to Curio’s cloud platform for secondary analysis, including optional imputation, reporting, and estimated breeding value (EBV) calculation.

    Questions worth asking when evaluating a genotyping platform

    For anyone specifying, validating, or procuring a genotyping method for a breeding program, a service lab, or a national database submission, a few questions surface the tradeoffs above concretely:

    • What is the panel’s ascertainment history, and how does expected call rate and minor allele frequency change for breeds or crossbreds outside the discovery population?

    • Is the genotype call an inference from signal intensity, or a direct sequence read? What happens at sites with adjacent variation? Is there a documented failure mode, and how is it detected?

    • What’s the redesign cycle time if a new marker needs to be added, and what does that cost?

    • Has the method, or the specific lab/pipeline, been validated against an orthogonal method? Is that concordance data available, including a breakdown of the discordant sites rather than just the aggregate rate?

    • For evaluation submission: is the method or lab certified by the relevant standards body (CDCB, ICAR, or equivalent), and under what panel design?

    Closing note

    Arrays, targeted sequencing, and low-pass sequencing with imputation are all legitimate, currently-deployed methods. But the accuracy conversation around arrays has largely stopped at the aggregate concordance number. The types of sites where arrays may perform inconsistently is not random and may be impactful to some programs. The accuracy advantage of NGS is worth weighing alongside other important considerations including cost, turnaround, and workflow fit when specifying a genotyping method for a breeding program or service pipeline.


    References:

    1. CDCB, “CDCB Keeps Pace with Advancing Genomic Technology Through Integration of Genotyping-by-Sequencing,” June 5, 2025: https://uscdcb.com/cdcb-keeps-pace-with-gbs

    2. Marina, et. al. (2021). Study on the concordance between different SNP-genotyping platforms in sheep. Animal Genetics. doi: 10.1111/age.13139

    3. Ren, et. al. (2024). Evaluating the efficacy of target capture sequencing for genotyping in cattle. Genes. doi: 10.3390/genes15091218

     

    For Research Use Only. Not for use in diagnostic procedures.