The massive growth of genomic data has made it necessary to have effective computational algorithms that are able to detect small, biologically significant sequence patterns which are called motifs. The Planted Motif Problem (PMP) that consists of locating (l, d) motif in a number of sequences with a maximum of d mutations is still computationally challenging because of its combinatorial complexity. PMM (Planted Motif Miner) is a hybrid approximate framework that we suggest to identify (l, d) planted motifs efficiently using pattern mining and deep learning. In this work, PMM has seven steps including data preprocessing, k-mer segmentation, frequent-pattern aggregation, motif extension, consensus motif mining, d-neighbor generation, and CNN-based classification module. The model uses multiprocessing and random sampling to deal with the exponential growth of d-neighbors with large (l, d). Experimental results on simulated and real data show that PMM are accurate ranging from (87% to 99%) based on length of (l, d) parameters and better than state of the art motif discovery algorithms, such as qPMS10, FMotif, GADEM, and MEME-ChIP. Its performance advantage is obvious on large instances where the current ways will normally fail or take over 24 hours to accomplish. In addition, the suggested algorithm can execute large (l, d) like (30, 10). The findings indicate that the suggested method is able to obtain > 99% accuracy and F1-scores above 97% with a minimum loss rate, when using moderate parameter configurations, such as (18, 6) and (26, 6). The sampling is necessary when the problem complexity is higher to make sure that it is feasible, the classification performance tends to decrease, but the CNN can still achieve strong accuracy and F1-scores over 82% in the most difficult settings.
V. V. S. Dileep, R. Navuduru, R. Gummadi, and P. Natarajan, “DNA sequencing using machine learning and deep learning algorithms,” International Journal of Innovative Technology and Exploring Engineering (IJITEE), vol. 11, no. 10, 2022.
D. Sharma and S. Rajasekaran, “A simple algorithm for (l, d) motif search,” Search, pp. 148-154, 2009.
C. Pizzi, “Motif discovery with compact approaches: Design and applications,” IntechOpen, 2011.
Q. Yu, H. Huo, X. Chen, H. Guo, J. S. Vitter, and J. Huan, “An efficient motif finding algorithm for large DNA data sets,” in IEEE International Conference on Bioinformatics and Biomedicine (BIBM), 2014.
J. Davila, S. Balla, and S. Rajasekaran, “Space and time efficient algorithms for planted motif search,” in V. N. Alexandrov et al., Eds., ICCS 2006, Part II, Lecture Notes in Computer Science, vol. 3992, pp. 822-829, Springer, 2006.
P. Xiao, S. Pal, and S. Rajasekaran, “qPMS10: A randomized algorithm for efficiently solving quorum planted motif search problem,” in IEEE International Conference on Bioinformatics and Biomedicine (BIBM), USA, 2016.
D. Quang and X. Xie, “DanQ: A hybrid convolutional and recurrent deep neural network for quantifying the function of DNA sequences,” Nucleic Acids Research, vol. 44, p. e107, 2016.
Z. Shen, W. Bao, and D.-S. Huang, “Recurrent neural network for predicting transcription factor binding sites,” Scientific Reports, vol. 8, pp. 1-10, 2018.
X. Pan, P. Rijnbeek, J. Yan, et al., “Prediction of RNA-protein sequence and structure binding preferences using deep convolutional and recurrent neural networks,” BMC Genomics, vol. 19, p. 511, 2018.
C. Angermueller, T. Pärnamaa, L. Parts, and O. Stegle, “Deep learning for computational biology,” Molecular Systems Biology, vol. 12, no. 7, p. 878, 2018.
S. Min, B. Lee, and S. Yoon, “Deep learning in bioinformatics,” Briefings in Bioinformatics, vol. 18, no. 5, pp. 851-869, 2017.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning, MIT Press, 2016.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, p. 436, 2015.
M. A. Di Gangi, S. Gaglio, C. La Bua, G. L. Bosco, and R. Rizzo, “A deep learning network for exploiting positional information in nucleosome-related sequences,” in International Conference on Bioinformatics and Biomedical Engineering, pp. 524-533, Springer, 2017.
D. Asir, S. Appavu, and E. Jebamalar, “Literature review on feature selection methods for high-dimensional data,” International Journal of Computer Applications, vol. 136, no. 1, pp. 9-17, 2016.
Y. He, Z. Shen, Q. Zhang, S. Wang, and D.-S. Huang, “A survey on deep learning in DNA/RNA motif mining,” Briefings in Bioinformatics, vol. 22, no. 4, pp. 1-10, 2021.
S. Wang and T. Huang, “Applications of deep learning in biomedicine,” in Reference Module in Biomedical Sciences, pp. 1-11, Elsevier, 2019.
Y. Li, C. Huang, L. Ding, et al., “Deep learning in bioinformatics: Introduction, application, and perspective in the big data era,” Methods, vol. 166, pp. 4-21, 2019.
R. Yamashita, M. Nishio, R. K. G. Do, and K. Togashi, “Convolutional neural networks: An overview and application in radiology,” Insights into Imaging, vol. 9, pp. 611-629, 2018.
N. Tajbakhsh, J. Y. Shin, S. R. Gurudu, et al., “Convolutional neural networks for medical image analysis: Full training or fine tuning,” IEEE Transactions on Medical Imaging, vol. 35, pp. 1299-1312, 2016.
Q. Li, W. Cai, X. Wang, et al., “Medical image classification with convolutional neural networks,” in 13th International Conference on Control Automation Robotics & Vision (ICARCV), pp. 844-848, 2014.
S. H. S. Basha, S. R. Dubey, V. Pulabaigari, and S. Mukherjee, “Impact of fully connected layers on performance of convolutional neural networks for image classification,” Neurocomputing, vol. 378, pp. 112-119, 2020.
C. Jia, M. B. Carson, Y. Wang, Y. Lin, and H. Lu, “A new exhaustive method and strategy for finding motifs in ChIP-enriched regions,” PLoS ONE, vol. 9, no. 1, p. e86044, 2014.
L. Li, “GADEM: A genetic algorithm guided formation of spaced dyads coupled with an EM algorithm for motif discovery,” Journal of Computational Biology, vol. 16, no. 2, pp. 317-329, 2009.
P. Machanick and T. L. Bailey, “MEME-ChIP: Motif analysis of large DNA datasets,” Bioinformatics, vol. 27, no. 12, pp. 1696-1697, 2011.