How Do Scientists Interpret Your Genome? What Actually Happens from Sequencing Reads to High-Confidence Variants
Sequencing: Converting biological molecules into digital reads. The first step of genome sequencing is fragmenting a physical sample (such as DNA from blood or saliva) and reading it with a sequencer to generate hundreds of millions of short DNA sequences, known as “sequencing reads”. There are currently two main sequencing technologies: Short-read sequencing (e.g., Illumina): Produces shorter (75–250 bases) but highly accurate reads. Due to its high throughput and cost-effectiveness, this is currently the most widely used foundational technology. Long-read sequencing (e.g., PacBio and Oxford Nanopore): Capable of reading ultra-long DNA fragments of thousands to millions of bases in a single pass.
What Is a “Pangenome”, and Why Is the Traditional Single Reference Genome No Longer Enough?
A pangenome is a novel, more representative genomic map that merges reference sequences from multiple individuals across diverse global populations—such as the diploid assemblies of 47 genetically diverse individuals combined by the Human Pangenome Reference Consortium—into a unified reference structure. Unlike traditional linear sequences, a pangenome is typically represented using graph data structures: nodes in the graph contain the set of known sequences within the population, while paths traversing these nodes compactly describe an individual’s unique DNA sequence. In this way, the pangenome embeds population genetic variation directly into the reference model.
Traditional single reference genomes (such as GRCh38) are no longer sufficient, primarily due to several core limitations:
How Much Did Accuracy Improve After DeepVariant Transformed DNA Data into Images?
After DeepVariant framed variant calling as an image classification problem, its accuracy achieved a major leap forward. Because a convolutional neural network (CNN) can process all relevant sequencing reads simultaneously as a multi-channel image, it successfully captured complex dependencies that traditional statistical algorithms (which mistakenly assumed read errors were independent) failed to handle.
According to the source data, the accuracy improvements were specifically reflected in the following dimensions:
Over 50% reduction in overall error rate: In the FDA Precision Medicine Truth Challenge, without needing to adjust parameters such as filtering thresholds, DeepVariant reduced total errors per genome by more than 50% compared to the second-ranked traditional algorithm (DeepVariant made only 4,652 errors, compared to 9,531 errors for the second-place algorithm).
Why Deep Learning Tools Like DeepVariant Are Better at “Reading” Genomes Than Traditional Algorithms
The fundamental reason deep learning tools like DeepVariant are better at “reading” genomes than traditional algorithms such as GATK is that they transform genomic variant calling from “expert-driven statistical modeling” into a “data-driven image classification problem.” This paradigm shift gives them several key advantages over traditional algorithms:
Breaking the traditional assumption of independent sequencing errors: When building statistical models, traditional algorithms (such as GATK using logistic regression and Hidden Markov Models) typically assume that errors in sequencing reads are independent of one another for mathematical tractability. In reality, however, systematic errors from sequencers and the complex structure of the genome itself lead to complex correlated errors.
Since a 21x sequencing depth can match the accuracy of traditional 30x, how much cost can this save for large-scale genomic sequencing?
According to the data, assuming a standard 30x whole-genome sequencing cost of $1,000 and a linear relationship between cost and sequencing depth, using DeepVariant can save approximately $250 per genome.
This is because at a sequencing depth of 22x to 23x, DeepVariant achieves the same total number of errors as traditional tools (such as GATK4-HC) at 30x depth.
Additionally, the data notes cost savings at higher depths: the error count of DeepVariant at 27x depth is roughly equivalent to that of GATK4-HC at 50x depth. Based on the same linear cost calculation, this translates to saving about $766 per genome under a 50x sequencing standard.
Can DeepVariant, a technology that transforms DNA into images, be used to detect complex somatic mutations such as cancer?
In fact, building on the core technology of DeepVariant, scientists specifically developed a new deep learning tool for detecting somatic mutations in cancer—DeepSomatic.
Detecting somatic mutations (such as those occurring in cancer) is more complex than detecting germline variants because cancer samples are typically mosaics composed of distinct clonal cell populations, accompanied by various complex mutational processes and signatures. DeepSomatic successfully applied image-based technology to this highly challenging field, with key features including:
Clever adaptation of image construction: DeepSomatic deeply adapted DeepVariant, with the most critical change being the modification of pileup images so that a single image contains aligned reads from both the “tumor sample” and the “normal control sample.”
How Ultra-Rapid Whole-Genome Sequencing Helps Diagnose Critical Rare Diseases in Hours—and Where Bottlenecks Remain
How does ultra-rapid whole-genome sequencing (WGS) achieve a diagnosis within hours?
In critical care settings like neonatal intensive care units, timely diagnosis of rare genetic diseases can save lives. By combining Nanopore long-read sequencing technology with deep learning, scientists have successfully compressed the time from whole-genome sequencing to diagnosis to under 8 hours (taking just 7 hours 18 minutes and 7 hours 48 minutes in cases involving a 57-year-old double-lung transplant candidate and a 14-month-old infant, respectively).
This content is for reading and understanding research reports. It does not constitute investment advice or trading signals.
Read in App
Read global research reports on mobile.
This content is for research reading and does not constitute investment advice.