Genome-Wide Association Study (GWAS)
Genome-Wide Association Study · Also known as: GWAS, genome-wide association analysis, whole-genome association study, WGAS
A genome-wide association study (GWAS) systematically tests hundreds of thousands to millions of single-nucleotide polymorphisms (SNPs) across the human genome for statistical association with a trait or disease. By comparing allele frequencies between cases and controls — or by regressing SNP genotypes on a quantitative phenotype — GWAS identifies genomic loci that harbor common genetic variants contributing to complex traits. Since its large-scale debut in 2007, GWAS has catalogued thousands of robust disease–variant associations across virtually every common human condition.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
+24 more
When to use it
Use GWAS when you have a well-phenotyped cohort (at minimum several thousand individuals; ideally tens of thousands or more) and want to discover genomic loci contributing to a complex trait or disease without a prior hypothesis about specific genes. It is the standard discovery tool for the common-variant architecture of polygenic traits. Do not use GWAS when your sample is too small (fewer than ~1,000 cases severely limits power for small-effect variants), when the trait is caused by rare high-penetrance variants (use rare-variant sequencing studies instead), when the phenotype is not standardised across samples, or when you need to establish causality directly (GWAS identifies association, not causation; Mendelian randomisation or functional follow-up is required for causal inference).
Strengths & limitations
- Hypothesis-free discovery across the entire genome, uncovering unexpected loci and biological pathways.
- Well-standardised pipeline with mature software (PLINK, REGENIE, BOLT-LMM) and reporting norms (GWAS Catalog).
- Results are highly reproducible and replicable when sample sizes are adequate and QC is rigorous.
- Enables downstream applications including polygenic risk scores, genetic correlation analysis, and Mendelian randomisation.
- Publicly available summary statistics allow meta-analysis and secondary analyses without access to individual-level data.
- Common variants identified by GWAS typically explain a small fraction of phenotypic variance individually; large samples are required to detect them reliably.
- GWAS identifies associated loci, not causal variants or genes — fine-mapping and functional studies are always required for mechanistic interpretation.
- Most GWAS have been conducted in populations of European ancestry, limiting generalisability and creating inequity in genomic medicine.
- Cannot detect rare variants (MAF < 1%) with standard array-based designs; whole-genome sequencing or targeted rare-variant studies are needed for the rare-variant spectrum.
- Population stratification and cryptic relatedness can produce spurious associations if not properly controlled.
Frequently asked
How large does my sample need to be?
Power depends on effect size, allele frequency, and desired significance threshold. For typical complex-trait GWAS (odds ratio ~1.1–1.2 per allele), you need tens of thousands of individuals to detect signals reliably. A rule of thumb is that at least 5,000–10,000 cases are needed for adequately powered discovery; biobank-scale analyses (100K–1M individuals) are now standard for polygenic traits with many small-effect loci. Use software such as GAS Power Calculator or PLINK to estimate power before study design.
What is the difference between GWAS and whole-genome sequencing studies?
GWAS typically uses SNP arrays that genotype common variants (MAF > 1%) at hundreds of thousands to a few million positions, with imputation extending coverage. Whole-genome sequencing (WGS) reads the entire genome and captures both common and rare variants. GWAS is cheaper and scales to very large N, making it powerful for common-variant discovery. WGS is needed to detect rare pathogenic variants (MAF < 0.1%) but currently requires smaller samples due to cost. Many studies now combine both approaches.
How do I identify the causal gene from a GWAS locus?
GWAS delivers a genomic address, not a gene. Causal gene identification requires fine-mapping (statistical tools such as FINEMAP or SuSiE to identify the most probable causal variant set), colocalization with eQTL data (e.g., GTEx) to link the signal to gene expression changes, and functional experiments (CRISPR, reporter assays, in vivo models). No purely computational analysis can substitute for experimental validation.
Can GWAS be used for quantitative traits, not just case-control studies?
Yes. GWAS uses linear regression for quantitative phenotypes (height, BMI, blood pressure, cognitive scores) and logistic regression for binary outcomes. Quantitative-trait GWAS is often more powerful per sample than case-control designs because every individual provides information proportional to their deviation from the mean, whereas in case-control studies the controls provide less variance. Mixed-model methods (BOLT-LMM, SAIGE, REGENIE) now handle both types efficiently even in the presence of relatedness.
What is a polygenic risk score and how does it relate to GWAS?
A polygenic risk score (PRS) aggregates the effects of many GWAS-identified SNPs — often hundreds of thousands — into a single numeric score for an individual. Each SNP contributes its effect size (beta or log-odds ratio) from GWAS summary statistics, weighted by the number of risk alleles the individual carries. PRS is a post-GWAS application used for risk stratification in clinical and research settings, not a method to discover associations.
Sources
- Wellcome Trust Case Control Consortium. (2007). Genome-wide association study of 14,000 cases of seven common diseases and 3,000 shared controls. Nature, 447(7145), 661–678. link ↗
- Visscher, P. M., Wray, N. R., Zhang, Q., Sklar, P., McCarthy, M. I., Brown, M. A., & Yang, J. (2017). 10 years of GWAS discovery: Biology, function, and translation. American Journal of Human Genetics, 101(1), 5–22. DOI: 10.1016/j.ajhg.2017.06.005 ↗
How to cite this page
ScholarGate. (2026, June 3). Genome-Wide Association Study. ScholarGate. https://scholargate.app/en/bioinformatics/genome-wide-association-study
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Copy Number Variation AnalysisBioinformatics↔ compare
- Epigenome-wide association studyBioinformatics↔ compare
- eQTL AnalysisBioinformatics↔ compare
- Pathway Enrichment AnalysisBioinformatics↔ compare
- RNA-seq Differential ExpressionBioinformatics↔ compare
- Variant CallingBioinformatics↔ compare