DSpace Repository

Prediction of Epigenetic Protein using Deep Learning and Protein Large Language Models

Show simple item record

dc.contributor.author Kaleem Ullah, 01-243241-010
dc.date.accessioned 2026-08-31T09:46:20Z
dc.date.available 2026-08-31T09:46:20Z
dc.date.issued 2026
dc.identifier.uri http://hdl.handle.net/123456789/21659
dc.description Supervised by Dr. Farman Ali en_US
dc.description.abstract Epigenetic proteins (EPs) are essential in regulating gene activity without altering the DNA sequence. Their role in numerous diseases, including cancer, neurodegenerative, and autoimmune disorders, has rendered them significant targets in drug discovery. It’s important to be able to accurately predict EPs because they play such a key role. Unfortunately, experimental methods for prediction are often costly and take a long time and required domain experts. In this study, we introduced a new deep learning framework that uses protein sequence data to make predictions about EPs. We obtained protein sequence data from Uniprot database and used two powerful pre-trained large language models, ProtBERT and UniREP, to turn protein sequences into rich numerical representations. Then, these embeddings are used as inputs for three different deep learning models: a Multi-Headed Ensemble Residual Convolutional Neural Network (MERCNN), a Bi-Directional Gated Recurrent Unit (Bi-GRU), and an Encoder-Decoder model. We systematically evaluated each model’s performance using both embedding models to determine the best combinations for predicting EP. Our model (MERCNN-EP) achieved an impressive 96.89% accuracy on the training dataset and 91.12% on the testing dataset, which was obtained from the UniProt database to evaluate the generalization ability of the proposed model. These results outperform previously published studies on epigenetic protein prediction, demonstrating the effectiveness and improved predictive capability of the proposed framework. Our results show that the MERCNN architecture works best when combined with ProtBERT embeddings. This research creates a strong and scalable framework for high-capacity EP screening, which could have a big impact on improving targeted therapeutic strategies in precision medicine. en_US
dc.language.iso en en_US
dc.publisher Computer Science en_US
dc.relation.ispartofseries MS (CS);T-4005
dc.subject Epigenetic Protein en_US
dc.subject Deep Learning en_US
dc.subject Large Language Models en_US
dc.title Prediction of Epigenetic Protein using Deep Learning and Protein Large Language Models en_US
dc.type Thesis en_US


Files in this item

This item appears in the following Collection(s)

Show simple item record

Search DSpace


Advanced Search

Browse

My Account