Network-based GWAS — Network-based Genome-Wide Association Study
Network-based Genome-Wide Association Study · Also known as: network GWAS, gene network GWAS, network-informed GWAS, NbGWAS
Network-based GWAS integrates conventional genome-wide association study results with biological network data — such as protein-protein interaction (PPI) networks or gene co-expression graphs — to identify disease-relevant gene modules or subnetworks. Instead of reporting only the top individual SNPs, this approach propagates association signals through molecular interaction networks, surfacing gene clusters whose collective signal implicates them in complex-trait biology even when no single variant reaches genome-wide significance alone.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use network-based GWAS when a standard GWAS has been performed but yields few or no genome-wide significant hits for a complex trait, and the goal is to recover biologically coherent signal from sub-threshold variants. It is particularly powerful for polygenic traits where many genes of small effect collectively drive disease risk. You need GWAS summary statistics (full p-value table, not just top hits), a reliable biological network relevant to the tissue or condition under study, and computational resources for permutation testing. Do not use it as a substitute for a well-powered primary GWAS — it amplifies existing signal; it cannot rescue a severely underpowered study. Avoid it when the primary GWAS is in a species or cell type for which no quality network exists, or when the research question requires variant-level resolution (e.g., fine-mapping a causal SNP) rather than gene-module discovery.
Strengths & limitations
- Recovers biologically meaningful signal from sub-threshold GWAS loci that would be discarded by conventional SNP-based thresholding.
- Produces interpretable gene-module outputs that connect directly to biological pathways and functional hypotheses.
- Reduces the multiple-testing burden relative to testing all genes individually by focusing on network-coherent modules.
- Agnostic to the specific network source — can be applied with PPI networks, co-expression networks, or pathway graphs depending on biological context.
- Leverages existing GWAS summary statistics, making it applicable as a post-hoc analysis without access to individual-level genotype data.
- Results depend heavily on the quality and completeness of the biological network; sparse or biased networks (over-representing well-studied genes) can skew module discovery.
- Gene-score aggregation from SNPs is a lossy compression step: LD structure and multi-SNP effects within a gene are only partially captured.
- Permutation-based significance testing is computationally intensive and may be infeasible for very large networks without high-performance computing.
- Identified modules reflect the topology of the input network; truly novel biology not yet represented in curated databases will be missed.
- Interpretation requires biological expertise to evaluate whether a significant module is mechanistically plausible rather than a network artefact.
Frequently asked
Do I need individual-level genotype data or are summary statistics sufficient?
Summary statistics (per-SNP p-values and effect sizes across all tested variants) are sufficient for most network-based GWAS tools. Individual-level genotype data are not required, making this approach applicable to publicly released GWAS summary data.
Which biological network should I use?
The choice should match the biological context of the trait. For most human complex-disease GWAS, STRING or BioGRID PPI networks are common starting points. For traits with strong tissue specificity (e.g., brain disorders, liver metabolism), tissue-specific co-expression networks derived from GTEx or TCGA are more appropriate. Using multiple networks and checking for consistent modules adds confidence.
How is this different from standard pathway enrichment analysis?
Pathway enrichment analysis tests whether predefined gene sets (e.g., GO terms, KEGG pathways) are over-represented among GWAS hits. Network-based GWAS is data-driven: it discovers modules de novo from the network topology without relying on predefined pathway annotations, and it uses the full distribution of association scores rather than a binary hit/non-hit classification.
What sample size GWAS is large enough to benefit from this approach?
There is no strict minimum, but the approach is most valuable when a GWAS has at least moderate power — typically tens of thousands of samples for complex traits. If the GWAS has very few participants and no variants approach suggestive significance, the gene-score signal is essentially noise and network propagation cannot recover it.
Can I apply this to non-human species?
Yes, provided a quality biological network exists for the species of interest. Model organism databases (e.g., FlyBase, WormBase, SGD) curate PPI networks for Drosophila, C. elegans, and yeast. For less-studied species, network data are sparse and results should be interpreted cautiously.
Sources
- Wang, Q., Yu, H., Zhao, Z., & Jia, P. (2015). EW_dmGWAS: edge-weighted dense module search for genome-wide association studies and gene expression profiles. Bioinformatics, 31(15), 2591–2594. link ↗
- Leiserson, M. D. M., Vandin, F., Wu, H.-T., Dobson, J. R., Eldridge, J. V., Thomas, J. L., Papoutsaki, A., Kim, Y., Niu, B., McLellan, M., Lawrence, M. S., Gonzalez-Perez, A., Tamborero, D., Cheng, Y., Ryslik, G. A., Lopez-Bigas, N., Getz, G., Ding, L., & Raphael, B. J. (2015). Pan-cancer network analysis identifies combinations of rare somatic mutations contributing to tumorigenesis. Nature Genetics, 47(2), 106–114. link ↗
How to cite this page
ScholarGate. (2026, June 3). Network-based Genome-Wide Association Study. ScholarGate. https://scholargate.app/en/bioinformatics/network-based-gwas
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Copy Number Variation AnalysisBioinformatics↔ compare
- eQTL AnalysisBioinformatics↔ compare
- Gene Set Enrichment AnalysisBioinformatics↔ compare
- Genome-wide association studyBioinformatics↔ compare
- Network-based eQTL analysisBioinformatics↔ compare
- Pathway Enrichment AnalysisBioinformatics↔ compare