Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Bioinformatics›Genome-Wide Association Study (GWAS)
Process / pipelineBioinformatics / omics

Genome-Wide Association Study (GWAS)

Genome-Wide Association Study · Also known as: GWAS, genome-wide association analysis, whole-genome association study, WGAS

A genome-wide association study (GWAS) systematically tests hundreds of thousands to millions of single-nucleotide polymorphisms (SNPs) across the human genome for statistical association with a trait or disease. By comparing allele frequencies between cases and controls — or by regressing SNP genotypes on a quantitative phenotype — GWAS identifies genomic loci that harbor common genetic variants contributing to complex traits. Since its large-scale debut in 2007, GWAS has catalogued thousands of robust disease–variant associations across virtually every common human condition.

ScholarGate
  1. Process / pipeline
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Genome-wide association study
Copy Number Variation An…Epigenome-wide associati…eQTL AnalysisPathway Enrichment Analy…RNA-seq Differential Exp…Variant CallingBayesian Copy Number Var…Bayesian epigenome-wide…Bayesian epigenome-wide…Bayesian eQTL analysis

+24 more

When to use it

Use GWAS when you have a well-phenotyped cohort (at minimum several thousand individuals; ideally tens of thousands or more) and want to discover genomic loci contributing to a complex trait or disease without a prior hypothesis about specific genes. It is the standard discovery tool for the common-variant architecture of polygenic traits. Do not use GWAS when your sample is too small (fewer than ~1,000 cases severely limits power for small-effect variants), when the trait is caused by rare high-penetrance variants (use rare-variant sequencing studies instead), when the phenotype is not standardised across samples, or when you need to establish causality directly (GWAS identifies association, not causation; Mendelian randomisation or functional follow-up is required for causal inference).

Strengths & limitations

Strengths
  • Hypothesis-free discovery across the entire genome, uncovering unexpected loci and biological pathways.
  • Well-standardised pipeline with mature software (PLINK, REGENIE, BOLT-LMM) and reporting norms (GWAS Catalog).
  • Results are highly reproducible and replicable when sample sizes are adequate and QC is rigorous.
  • Enables downstream applications including polygenic risk scores, genetic correlation analysis, and Mendelian randomisation.
  • Publicly available summary statistics allow meta-analysis and secondary analyses without access to individual-level data.
Limitations
  • Common variants identified by GWAS typically explain a small fraction of phenotypic variance individually; large samples are required to detect them reliably.
  • GWAS identifies associated loci, not causal variants or genes — fine-mapping and functional studies are always required for mechanistic interpretation.
  • Most GWAS have been conducted in populations of European ancestry, limiting generalisability and creating inequity in genomic medicine.
  • Cannot detect rare variants (MAF < 1%) with standard array-based designs; whole-genome sequencing or targeted rare-variant studies are needed for the rare-variant spectrum.
  • Population stratification and cryptic relatedness can produce spurious associations if not properly controlled.

Frequently asked

How large does my sample need to be?

Power depends on effect size, allele frequency, and desired significance threshold. For typical complex-trait GWAS (odds ratio ~1.1–1.2 per allele), you need tens of thousands of individuals to detect signals reliably. A rule of thumb is that at least 5,000–10,000 cases are needed for adequately powered discovery; biobank-scale analyses (100K–1M individuals) are now standard for polygenic traits with many small-effect loci. Use software such as GAS Power Calculator or PLINK to estimate power before study design.

What is the difference between GWAS and whole-genome sequencing studies?

GWAS typically uses SNP arrays that genotype common variants (MAF > 1%) at hundreds of thousands to a few million positions, with imputation extending coverage. Whole-genome sequencing (WGS) reads the entire genome and captures both common and rare variants. GWAS is cheaper and scales to very large N, making it powerful for common-variant discovery. WGS is needed to detect rare pathogenic variants (MAF < 0.1%) but currently requires smaller samples due to cost. Many studies now combine both approaches.

How do I identify the causal gene from a GWAS locus?

GWAS delivers a genomic address, not a gene. Causal gene identification requires fine-mapping (statistical tools such as FINEMAP or SuSiE to identify the most probable causal variant set), colocalization with eQTL data (e.g., GTEx) to link the signal to gene expression changes, and functional experiments (CRISPR, reporter assays, in vivo models). No purely computational analysis can substitute for experimental validation.

Can GWAS be used for quantitative traits, not just case-control studies?

Yes. GWAS uses linear regression for quantitative phenotypes (height, BMI, blood pressure, cognitive scores) and logistic regression for binary outcomes. Quantitative-trait GWAS is often more powerful per sample than case-control designs because every individual provides information proportional to their deviation from the mean, whereas in case-control studies the controls provide less variance. Mixed-model methods (BOLT-LMM, SAIGE, REGENIE) now handle both types efficiently even in the presence of relatedness.

What is a polygenic risk score and how does it relate to GWAS?

A polygenic risk score (PRS) aggregates the effects of many GWAS-identified SNPs — often hundreds of thousands — into a single numeric score for an individual. Each SNP contributes its effect size (beta or log-odds ratio) from GWAS summary statistics, weighted by the number of risk alleles the individual carries. PRS is a post-GWAS application used for risk stratification in clinical and research settings, not a method to discover associations.

Sources

  1. Wellcome Trust Case Control Consortium. (2007). Genome-wide association study of 14,000 cases of seven common diseases and 3,000 shared controls. Nature, 447(7145), 661–678. link ↗
  2. Visscher, P. M., Wray, N. R., Zhang, Q., Sklar, P., McCarthy, M. I., Brown, M. A., & Yang, J. (2017). 10 years of GWAS discovery: Biology, function, and translation. American Journal of Human Genetics, 101(1), 5–22. DOI: 10.1016/j.ajhg.2017.06.005 ↗

How to cite this page

ScholarGate. (2026, June 3). Genome-Wide Association Study. ScholarGate. https://scholargate.app/en/bioinformatics/genome-wide-association-study

Related methods

Copy Number Variation AnalysisEpigenome-wide association studyeQTL AnalysisPathway Enrichment AnalysisRNA-seq Differential ExpressionVariant Calling

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Copy Number Variation AnalysisBioinformatics↔ compare
  • Epigenome-wide association studyBioinformatics↔ compare
  • eQTL AnalysisBioinformatics↔ compare
  • Pathway Enrichment AnalysisBioinformatics↔ compare
  • RNA-seq Differential ExpressionBioinformatics↔ compare
  • Variant CallingBioinformatics↔ compare
Compare side by side →

Referenced by

Bayesian Copy Number Variation AnalysisBayesian epigenome-wide association studyBayesian epigenome-wide association study in educational researchBayesian eQTL analysisBayesian GWASCopy Number Variation AnalysisDifferential Copy Number Variation AnalysisDifferential Epigenome-Wide Association StudyDifferential eQTL AnalysisEpigenome-wide association studyEpigenome-wide association study in educational researcheQTL AnalysisMachine learning-assisted copy number variation analysisMachine learning-assisted epigenome-wide association studyMachine learning-assisted expression quantitative trait loci analysisMachine learning-assisted genome-wide association studyMachine learning-assisted phylogenetic analysisMulti-omics epigenome-wide association studyMulti-omics eQTL analysisNetwork-based epigenome-wide association studyNetwork-based eQTL analysisNetwork-based GWASNetwork-based Phylogenetic AnalysisNetwork-based variant callingPhylogenetic AnalysisSequence AlignmentSingle-cell eQTL analysisSingle-cell GWASTime-series copy number variation analysisTime-series eQTL analysisTime-series phylogenetic analysisVariant Calling

Similar methods

Machine learning-assisted genome-wide association studyNetwork-based GWASSingle-cell GWASPolygenic Risk ScoreBayesian GWASGenome-wide association study in educational researcheQTL AnalysisEpigenome-wide association study

Related reference concepts

Genome-Wide Association Studies and Variant DiscoveryGWAS Design, Execution, and Statistical MethodsGenetic Basis of Disease SusceptibilityGenetic Basis of Complex DiseaseMissing Heritability and Polygenic ArchitecturePopulation Stratification and Ancestry in GWAS

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Genome-wide association study (Genome-Wide Association Study). Retrieved 2026-07-20 from https://scholargate.app/en/bioinformatics/genome-wide-association-study · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Klein et al. (age-related macular degeneration GWAS, 2005); landmark scale: Wellcome Trust Case Control Consortium (2007)
Year
2005–2007
Type
Observational genomic association study
DataType
SNP genotype array data or whole-genome sequencing data; binary or quantitative phenotype
Subfamily
Bioinformatics / omics
Related methods
Copy Number Variation AnalysisEpigenome-wide association studyeQTL AnalysisPathway Enrichment AnalysisRNA-seq Differential ExpressionVariant Calling
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account