Inferring binding specificities of human transcription factors with the wisdom of crowds [preprint]
Gryzunov, Nikita ; Penzar, Dmitry ; Kamenets, Vasilii ; Vyaltsev, Valery ; Kozin, Ivan ; Eliseeva, Irina A ; Nozdrin, Vladimir ; Vorontsov, Ilya E ; Bushuev, Sergey ; Strekalovskikh, Vadim ... show 10 more
Authors
Penzar, Dmitry
Kamenets, Vasilii
Vyaltsev, Valery
Kozin, Ivan
Eliseeva, Irina A
Nozdrin, Vladimir
Vorontsov, Ilya E
Bushuev, Sergey
Strekalovskikh, Vadim
Zinkevich, Arsenii
Andrews, Gregory
Bedarew, Matwej
Blass, Ido
Frolov, Dmitry
Lariushina, Iuliia
Moore, Jill
Orenstein, Yaron
Roev, German
Salimov, Danil
Shimshoviz, Noam
Tziony, Ido
Weng, Zhiping
Bucher, Philipp
Deplancke, Bart
Fornes, Oriol
Grau, Jan
Grosse, Ivo
Jolma, Arttu
Kolpakov, Fedor A
Makeev, Vsevolod J
Hughes, Timothy R
Kulakovskiy, Ivan V
Student Authors
Faculty Advisor
Academic Program
UMass Chan Affiliations
Document Type
Publication Date
Subject Area
Files
Embargo Expiration Date
Link to Full Text
Abstract
DNA motif discovery and, particularly, computational modeling of transcription factor binding motifs, has been a mecca of algorithmic bioinformatics for several decades. Here, we report the results of the largest open community challenge in Inferring BInding Specificities (IBIS), where participants all over the world were invited to construct binding specificity models from multi-assay experimental data for poorly studied human transcription factors. The submissions were rigorously tested against a rich held-out dataset. Benchmarking demonstrated a consistent advantage of properly designed deep learning models over traditional positional weight matrices and other machine learning methods. Yet, the positional weight matrices displayed a surprisingly strong performance out of the box, being only slightly behind the best deep learning models. A post-challenge assessment of a selection of other deep learning methods further solidified this finding. IBIS highlights the power of benchmarking in finding adequate DNA motif representations, emphasizes the pros and cons of various machine learning methods applied to DNA motif modeling, and establishes a rich dataset, benchmarking protocols, and computational framework for a fair cross-platform evaluation of future models of transcription factor binding motifs in DNA sequences.
Source
Gryzunov N, Penzar D, Kamenets V, Vyaltsev V, Kozin I, Eliseeva IA, Nozdrin V, Vorontsov IE, Bushuev S, Strekalovskikh V, Zinkevich A, Andrews G, Bedarew M, Blass I, Frolov D, Lariushina I, Moore J, Orenstein Y, Roev G, Salimov D, Shimshoviz N, Tziony I, Weng Z; IBIS Consortium; GRECO-BIT/Codebook Consortium; Bucher P, Deplancke B, Fornes O, Grau J, Grosse I, Jolma A, Kolpakov FA, Makeev VJ, Hughes TR, Kulakovskiy IV. Inferring binding specificities of human transcription factors with the wisdom of crowds. bioRxiv [Preprint]. 2025 Nov 17:2025.11.16.688692. doi: 10.1101/2025.11.16.688692. PMID: 41332515; PMCID: PMC12667932.
Year of Medical School at Time of Visit
Sponsors
Dates of Travel
DOI
Permanent Link to this Item
PubMed ID
Other Identifiers
Notes
This article is a preprint. Preprints are preliminary reports of work that have not been certified by peer review.