De Novo Transcriptome Assembly
De Novo RNA-Seq Transcriptome Assembly · Also known as: transcriptome assembly, de novo assembly, RNA-Seq assembly
De novo transcriptome assembly reconstructs full-length messenger RNA sequences directly from sequencing reads without requiring a reference genome. Pioneered by Regev, Haas, and colleagues, this pipeline enables transcript discovery in non-model organisms and detection of novel isoforms, fusion genes, and splice variants.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use de novo assembly to discover transcripts in non-model organisms, identify splice variants, and detect gene fusions. It is valuable when reference genomes are unavailable or incomplete. Avoid de novo assembly for expression quantification alone; reference-guided methods are faster and more accurate for known genes.
Strengths & limitations
- Enables transcript discovery without a reference genome
- Detects novel isoforms, fusion genes, and non-standard splice events
- Scalable to large transcriptomes with adequate sequencing depth
- Unbiased approach reveals unexpected transcribed regions
- De novo assembly is computationally intensive; requires substantial computing resources
- Assembly quality depends critically on sequencing depth and read length
- Distinguishing true isoforms from sequencing artifacts is challenging
- Low-abundance transcripts are often lost or misassembled
Frequently asked
How much sequencing coverage do I need for reliable de novo assembly?
Recommended coverage is 20-50x for complex transcriptomes; 10-20x for simpler genomes. Higher coverage improves isoform recovery and reduces misassembly. Coverage should be assessed per transcript: highly expressed genes assemble well at <5x, while rare transcripts require >50x for reliable detection.
How do I distinguish genuine novel isoforms from misassembled artifacts?
Validate using independent datasets, long-read sequencing, or RT-PCR. Check for consistent exon-exon junctions across samples. Novel isoforms with single-sample support are likely artifacts. Abundance estimates help; isoforms with very low read support across replicates are suspicious.
Can de novo assembly be used for genome-wide expression quantification?
Not reliably. Quantification from de novo assemblies is biased toward highly assembled transcripts and is sensitive to misassembly errors. For expression quantification, use reference-guided mapping to known gene models or quantify directly against reference genomes.
Sources
- Grabherr, M. G., Haas, B. J., Yassour, M., Levin, J. Z., Thompson, D. A., Amit, I., ... & Regev, A. (2011). Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nature Biotechnology, 29(7), 644-652. DOI: 10.1038/nbt.1883 ↗
- Haas, B. J., Papanicolaou, A., Yassour, M., Grabherr, M., Blood, P. D., Bowden, J., ... & Regev, A. (2013). De novo transcript sequence reconstruction from RNA-seq using the Trinity platform for reference generation and analysis. Nature Protocols, 8(8), 1494-1512. DOI: 10.1038/nprot.2013.084 ↗
- Pertea, M., Pertea, G. M., Antonescu, C. M., Chang, T. C., Mendell, J. T., & Salzberg, S. L. (2015). StringTie enables improved assembly of novel transcripts from RNA-seq data. Nature Biotechnology, 33(3), 290-295. link ↗
How to cite this page
ScholarGate. (2026, June 3). De Novo RNA-Seq Transcriptome Assembly. ScholarGate. https://scholargate.app/en/bioinformatics/de-novo-transcriptome-assembly
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- CRISPR Screen AnalysisBioinformatics↔ compare
- HMMER Profile SearchBioinformatics↔ compare
- Metagenomic BinningBioinformatics↔ compare