Computational and statistical considerations in the analysis of metagenomic data (2nd edition)
Book chapter, 2024

In shotgun metagenomics, microbial communities are studied by sequencing DNA fragments randomly sampled directly from environmental and clinical samples. The resulting data is massive, potentially consisting of billions of sequence reads describing millions of microbial genes. The interpretation of metagenomic data is therefore nontrivial and dependent on dedicated computational and statistical methods. In this chapter, we discuss the many challenges associated with the analysis of shotgun metagenomic data. First, we address computational issues related to the quantification of genes in metagenomes. We describe algorithms for efficient sequence comparisons, recommended practices for setting up data workflows, and modern high-performance computing resources that can be used to perform the analysis. Next, we outline the statistical aspects, including removal of systematic errors and identification of differences between microbial communities from different experimental conditions. We conclude by underlining the increasing importance of efficient and reliable computational and statistical solutions in the analysis of large metagenomic datasets.

Author

Fredrik Boulund

Karolinska Institutet

Mariana Buongermino Pereira

University of Gothenburg

Chalmers, Mathematical Sciences, Applied Mathematics and Statistics

Viktor Jonsson

Chalmers, Life Sciences, Systems and Synthetic Biology

University of Gothenburg

Erik Kristiansson

University of Gothenburg

Chalmers, Mathematical Sciences, Applied Mathematics and Statistics

Metagenomics (Second Edition)

83-104
9780081022689 (ISBN)

Subject Categories (SSIF 2025)

Bioinformatics (Computational Biology)

Bioinformatics and Computational Biology

Microbiology

DOI

10.1016/B978-0-323-91631-8.00001-9

More information

Latest update

7/30/2026