Learning the Regulatory Code of Gene Expression

Jan Zrimec; Filip Buric; Mariia Kokina; Victor Garcia; Aleksej Zelezniak

doi:10.3389/fmolb.2021.673363

Learning the Regulatory Code of Gene Expression
Reviewartikel, 2021

Data-driven machine learning is the method of choice for predicting molecular phenotypes from nucleotide sequence, modeling gene expression events including protein-DNA binding, chromatin states as well as mRNA and protein levels. Deep neural networks automatically learn informative sequence representations and interpreting them enables us to improve our understanding of the regulatory code governing gene expression. Here, we review the latest developments that apply shallow or deep learning to quantify molecular phenotypes and decode the cis-regulatory grammar from prokaryotic and eukaryotic sequencing data. Our approach is to build from the ground up, first focusing on the initiating protein-DNA interactions, then specific coding and non-coding regions, and finally on advances that combine multiple parts of the gene and mRNA regulatory structures, achieving unprecedented performance. We thus provide a quantitative view of gene expression regulation from nucleotide sequence, concluding with an information-centric overview of the central dogma of molecular biology.

machine learning

chromatin accessibility

mRNA & protein abundance

gene regulatory structure

cis-regulatory grammar

gene expression prediction

deep neural networks

regulatory genomics

Författare

Jan Zrimec

Chalmers, Biologi och bioteknik, Systembiologi

Forskning Andra publikationer

Filip Buric

Chalmers, Biologi och bioteknik, Systembiologi

Forskning Andra publikationer

Mariia Kokina

Chalmers, Biologi och bioteknik, Systembiologi

Danmarks Tekniske Universitet (DTU)

Forskning Andra publikationer

Victor Garcia

Zürcher Hochschule für Angewandte Wissenschaften

Aleksej Zelezniak

Chalmers, Biologi och bioteknik, Systembiologi

SciLifeLab

Forskning Andra publikationer

Frontiers in Molecular Biosciences

2296889X (eISSN)

Vol. 8 673363

Ämneskategorier (SSIF 2011)

Bioinformatik (beräkningsbiologi)

Bioinformatik och systembiologi

Genetik

DOI

10.3389/fmolb.2021.673363

Publikationsdata kopplat till DOI

PubMed

34179082

Mer information

Senast uppdaterat

2021-12-08

Learning the Regulatory Code of Gene Expression Reviewartikel, 2021

Författare

Jan Zrimec

Filip Buric

Mariia Kokina

Victor Garcia

Aleksej Zelezniak

Frontiers in Molecular Biosciences

Ämneskategorier (SSIF 2011)

DOI

PubMed

Mer information

Senast uppdaterat

Learning the Regulatory Code of Gene Expression
Reviewartikel, 2021