Deep learning as a tool to better understand transcription factor binding across cell types and species
Andrews, Gregory
Citations
Authors
Student Authors
Faculty Advisor
Academic Program
UMass Chan Affiliations
Document Type
Publication Date
Subject Area
Collections
Files
Embargo Expiration Date
Link to Full Text
Abstract
Deep learning has transformed our everyday lives. Facebook facial recognition, Netflix personalized recommendations and ChatGPT are all powered by deep learning. Neural networks, inspired by human intelligence, are the workhorse of deep learning, and are capable of learning complex relationships and patterns within large amounts of heterogeneous data. However, they are often referred to as black boxes due to the complex nature of these patterns and the difficulty involved in discerning them. This work seeks to open this metaphorical black box and leverage the information learned by these networks to better understand transcription factor (TF) binding. Convolutional neural networks (CNNs) excel at learning patterns from images. Nucleotide sequences are simply 1D images, a fact we use to develop a CNN-based motif discovery algorithm that outperforms classical and other deep learning-based approaches. We use our method and thousands of publicly available TF ChIP-seq experiments to annotate the binding sites of 367 human TFs. We then investigated the evolutionary conservation of these TF binding sites (TFBSs) in the mammalian lineage. We next demonstrate CNNs can be used to predict TF binding across cell types using only sequence and chromatin accessibility data. Lastly, we highlight the application of CNNs in unsupervised learning, specifically in the context of clustering brain specific regulatory elements based on sequence features. Altogether, the results presented herein highlight the importance of carefully constructing and training CNNs to achieve state of the art performance and gain the most biologically meaningful insights when trained on regulatory genomic data.