Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Bioinformatics›Variant Calling — Genomic Variant Calling
Process / pipelineBioinformatics / omics

Variant Calling — Genomic Variant Calling

Genomic Variant Calling · Also known as: SNP calling, genotyping from sequencing, mutation detection, variant detection

Variant calling is the computational process of identifying positions in a sequenced genome that differ from a reference sequence — including single nucleotide polymorphisms (SNPs), small insertions and deletions (indels), and structural variants. It transforms aligned sequencing reads into an interpretable catalogue of genetic differences, forming the foundation for population genetics, disease-gene discovery, and clinical genomics applications.

ScholarGate
  1. Process / pipeline
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Variant Calling
Copy Number Variation An…Epigenome-wide associati…Genome-wide association…RNA-seq Differential Exp…Sequence AlignmentSingle-cell variant call…Bayesian ChIP-seq peak c…Bayesian Copy Number Var…Bayesian Proteomics Anal…Bayesian RNA-seq differe…

+14 more

When to use it

Use variant calling when you have whole-genome, whole-exome, or targeted sequencing data and need to identify germline (inherited) or somatic (acquired) genetic variants. It is the standard first step in disease-gene mapping, pharmacogenomics, population diversity studies, and cancer genomics. Do not apply standard germline variant callers to single-cell sequencing data without specialist tools, as dropout and amplification noise require different statistical treatment. Avoid the method when coverage is very low (mean depth below approximately 10×) because insufficient read evidence leads to high false-negative rates and unreliable genotype calls.

Strengths & limitations

Strengths
  • Provides a genome-wide, base-resolution catalogue of genetic differences in a single experiment.
  • Well-supported by mature, widely adopted tools (GATK, SAMtools/bcftools, DeepVariant) with extensive documentation and community support.
  • Applicable to both germline and somatic variant detection with appropriate caller selection.
  • Scales readily to cohort-level analyses (joint genotyping across hundreds to thousands of samples) enabling population-level studies.
  • Standard VCF output integrates seamlessly with downstream association, annotation, and visualisation workflows.
Limitations
  • Accuracy is highly dependent on read coverage depth; low-coverage regions produce unreliable or absent calls.
  • Difficult genomic regions — repeats, segmental duplications, and low-complexity sequences — are systematically under-called by short-read technologies.
  • Somatic variant calling in tumour samples requires matched normal tissue and specialist tools; standard germline pipelines produce high false-positive rates on tumour data.
  • Structural variants (large deletions, inversions, translocations) are poorly captured by standard SNP/indel callers and require dedicated long-read or paired-end algorithms.
  • Computational resource requirements (storage, CPU, memory) are substantial for large cohorts or whole-genome data.

Frequently asked

What coverage depth do I need for reliable variant calling?

For germline SNP/indel calling with standard tools, a mean depth of 30× is widely recommended for high sensitivity and specificity. Many exome studies use 50–100× to compensate for uneven capture efficiency. Low-pass whole-genome sequencing (1–10×) is possible with specialised imputation-based approaches but is not suitable for rare-variant discovery in individual samples.

What is the difference between GATK HaplotypeCaller and SAMtools/bcftools?

HaplotypeCaller performs local de-novo assembly of haplotypes around candidate variant sites before calling, which makes it more accurate in regions with multiple nearby variants (e.g., complex indels) but more computationally intensive. SAMtools/bcftools uses a faster pileup-based Bayesian model that is effective for high-coverage data. For human germline calling, GATK is generally preferred; for organisms with less well-characterised variation or for rapid survey work, bcftools is a practical alternative.

When should I use VQSR versus hard filtering?

VQSR is recommended when you have a large enough call set (typically at least ~30 whole-genome samples or a well-powered exome cohort) and a validated truth set for the organism. It provides better-calibrated trade-offs between sensitivity and specificity than arbitrary threshold cuts. Hard filtering using recommended quality metrics is appropriate for smaller studies, non-human organisms lacking validated truth sets, or targeted panel data.

Can I call variants from RNA-seq data?

Yes, with caveats. GATK provides a documented best-practices workflow for RNA-seq variant calling that includes a splice-aware alignment step (STAR two-pass) and additional filtering for RNA-specific artefacts such as splicing-induced false positives. RNA-seq variant calling is useful for detecting variants expressed in a tissue of interest, but it covers only expressed regions and is not a substitute for DNA sequencing for comprehensive variant discovery.

How is single-sample calling different from joint genotyping?

In single-sample calling, each sample is processed independently and the resulting VCFs are merged. In joint genotyping (e.g., GATK GenotypeGVCFs), all samples are called together, allowing the model to use population-level allele frequency evidence to improve genotyping accuracy, particularly for rare variants. Joint genotyping is preferred for cohort studies because it reduces false negatives for rare alleles observed in only one or a few samples.

Sources

  1. McKenna, A., Hanna, M., Banks, E., Sivachenko, A., Cibulskis, K., Kernytsky, A., ... & DePristo, M. A. (2010). The Genome Analysis Toolkit: A MapReduce framework for analyzing next-generation DNA sequencing data. Genome Research, 20(9), 1297–1303. DOI: 10.1101/gr.107524.110 ↗
  2. Li, H., Handsaker, B., Wysoker, A., Fennell, T., Ruan, J., Homer, N., ... & Durbin, R. (2009). The Sequence Alignment/Map format and SAMtools. Bioinformatics, 25(16), 2078–2079. DOI: 10.1093/bioinformatics/btp352 ↗

How to cite this page

ScholarGate. (2026, June 3). Genomic Variant Calling. ScholarGate. https://scholargate.app/en/bioinformatics/variant-calling

Related methods

Copy Number Variation AnalysisEpigenome-wide association studyGenome-wide association studyRNA-seq Differential ExpressionSequence AlignmentSingle-cell variant calling

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Copy Number Variation AnalysisBioinformatics↔ compare
  • Epigenome-wide association studyBioinformatics↔ compare
  • Genome-wide association studyBioinformatics↔ compare
  • RNA-seq Differential ExpressionBioinformatics↔ compare
  • Sequence AlignmentBioinformatics↔ compare
  • Single-cell variant callingBioinformatics↔ compare
Compare side by side →

Referenced by

Bayesian ChIP-seq peak callingBayesian Copy Number Variation AnalysisBayesian Proteomics AnalysisBayesian RNA-seq differential expressionBayesian Sequence AlignmentBayesian Variant CallingChIP-seq Peak CallingCopy Number Variation AnalysisDifferential ChIP-seq peak callingGenome-wide association studyMachine learning-assisted ChIP-seq peak callingMachine learning-assisted copy number variation analysisNetwork-based copy number variation analysisNetwork-based Phylogenetic AnalysisNetwork-based variant callingPhylogenetic AnalysisRNA-seq Differential ExpressionSequence AlignmentSingle-cell Copy Number Variation AnalysisSingle-cell Phylogenetic AnalysisTime-series copy number variation analysisTime-series phylogenetic analysis

Similar methods

Bayesian Variant CallingMachine learning-assisted variant callingSingle-cell variant callingDifferential Variant CallingCopy Number Variation AnalysisNetwork-based variant callingMachine learning-assisted copy number variation analysisGenome-wide association study

Related reference concepts

Nucleotide Diversity and Variant ClassificationFunctional Annotation of Genomic VariantsCopy Number Variants: Detection and ClassificationGenome Variation and Population GenomicsQuality Control and Error Correction in SequencingGenetic Mutation Analysis and Interpretation

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Variant Calling (Genomic Variant Calling). Retrieved 2026-07-20 from https://scholargate.app/en/bioinformatics/variant-calling · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Li et al. (SAMtools/bcftools, 2009); McKenna et al. (GATK, 2010)
Year
2009–2010 (modern high-throughput era)
Type
Computational genomics pipeline
DataType
Aligned next-generation sequencing reads (BAM/CRAM files)
Subfamily
Bioinformatics / omics
Related methods
Copy Number Variation AnalysisEpigenome-wide association studyGenome-wide association studyRNA-seq Differential ExpressionSequence AlignmentSingle-cell variant calling
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account