<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
<title>Department of Computer Sciences (BUIC-E-8)</title>
<link href="http://hdl.handle.net/123456789/13167" rel="alternate"/>
<subtitle/>
<id>http://hdl.handle.net/123456789/13167</id>
<updated>2026-09-09T16:06:59Z</updated>
<dc:date>2026-09-09T16:06:59Z</dc:date>
<entry>
<title>Data Driven Earthquake Prediction using Intelligent Techniques</title>
<link href="http://hdl.handle.net/123456789/21658" rel="alternate"/>
<author>
<name>Muhammad Ibrahim Lodhi, 01-249241-007</name>
</author>
<id>http://hdl.handle.net/123456789/21658</id>
<updated>2026-09-01T05:08:22Z</updated>
<published>2026-01-01T00:00:00Z</published>
<summary type="text">Data Driven Earthquake Prediction using Intelligent Techniques
Muhammad Ibrahim Lodhi, 01-249241-007
Earthquakes represent one of the most devastating natural hazards, claiming thousands of lives and causing trillions in economic damage annually, particularly in seismically active regions like the Himalayan arc and Pacific Ring of Fire . Despite advances in seismic monitoring, accurate prediction of earthquake magnitude the key determinant of potential destruction remains elusive due to the nonlinear, chaotic dynamics of fault systems and the rarity of large events in historical catalogs. In conventional probabilistic seismic hazard analysis, long term risk predictions can be made, but reliable short term predictions are not possible, which makes society prone to abrupt ruptures. This thesis works to fill this crucial research gap by proposing and developing a new deep learning model to predict magnitude based on multi modal precursor. Although there have been some recent attempts to apply machine learning to seismicity forecasting, there have been several models still using simple regression models or simple neural networks, which are not capable of handling spatiotemporal correlations effectively. Although models such as random forest approaches have had moderate success at the regional scale, they also tend to neglect long term temporal correlations and lack generability across different tectonic settings. Although there have been some promising approaches using transformers for time series data, there have been no attempts to apply these models to seismicity forecasting, mainly due to their high complexity and the need for specific adaptations. These models, although successful, also emphasize the need to develop models that can integrate heterogeneous features of seismicity, such as b value anomaly, ionospheric TEC, and waveform magnitude, effectively, which can also provide probabilistic forecasting. In light of these challenges, this thesis develops the approach for earthquake magnitude prediction using modern deep learning and making a particular focus on: revealing temporal dynamics and predictive uncertainty. It focuses on an enhanced Temporal Fusion Transformer (TFT) architecture combining bidirectional LSTMs, GRUs, multi head attention, and gating mechanisms that combine past observations with simple future covariates and static spatial contextual information into a single unified forecasting model. To overcome the small size and imbalance in conventional seismic catalogs, this study constructs an augmented dataset with a model based synthetic data generator and relies on the assumption that the added records preserve the statistical and spatial properties of the original events. The original data and augmented are subjected to extensive feature engineering comprising lagged magnitudes, rolling statistics, interv action terms, and cluster based spatial descriptors, followed by their transformation into supervised sequences to model time series. The TFT model is trained using a quantile loss function to produce probabilistic forecasts at multiple quantiles, allowing the derivation of both point estimates and prediction intervals for earthquake magnitude. The training procedure incorporates regularization techniques such as dropout, weight decay, and gradient clipping, together with adaptive learning rate scheduling and validation based check pointing to promote stable convergence and generalization. Model performance is evaluated on a temporally held out test set using root mean squared error (RMSE), mean absolute error (MAE), and average quantile loss, alongside visual diagnostics such as prediction truth scatter plots, residual analysis, and time series overlays. In addition, Monte Carlo dropout is employed at inference time to form an ensemble of stochastic predictions, providing an estimate of epistemic uncertainty and enabling a comparison between single model and ensemble behavior. The results indicate that the proposed TFT architecture, combined with carefully designed augmentation and feature engineering, can achieve low prediction errors while generating well calibrated prediction intervals for normalized earthquake magnitudes. Spatial comparison plots show that the synthetic events closely follow the geographic distribution of the original Catalog, suggesting that the augmented dataset is suitable for training without distorting the underlying seismic patterns. Overall, the thesis demonstrates that attention based sequence models, when paired with realistic data augmentation and rigorous evaluation, offer a promising direction for probabilistic earthquake magnitude forecasting and for quantifying the uncertainty inherent in such predictions.
Supervised by Dr. Saba Mahmood
</summary>
<dc:date>2026-01-01T00:00:00Z</dc:date>
</entry>
<entry>
<title>Prediction of Tumor-Homing Peptide by Fusing Deep Learning with Protein Language Models</title>
<link href="http://hdl.handle.net/123456789/21660" rel="alternate"/>
<author>
<name>Waqas Ghafoor, 01-243232-011</name>
</author>
<id>http://hdl.handle.net/123456789/21660</id>
<updated>2026-08-31T09:54:43Z</updated>
<published>2026-01-01T00:00:00Z</published>
<summary type="text">Prediction of Tumor-Homing Peptide by Fusing Deep Learning with Protein Language Models
Waqas Ghafoor, 01-243232-011
Tumor-homing peptides (THPs) are short amino acid sequences that minimize off-target toxicity while selectively binding to molecular markers on tumor cells to enable targeted drug delivery and diagnostic imaging. The development of peptide-based cancer treatments depends on the accurate identification of THPs, but conventional experimental screening is time-consuming and previous computational predictors have shown limitations. For example, early machine-learning models (e.g., SVM-based approaches by Sharma et al. 2013) only achieved moderate accuracy because of small datasets and limited feature scope. These models relied on basic sequence features like amino acid composition. More potent tools for THP prediction have been made available by recent developments in artificial intelligence. As demonstrated by models like PLMTHP and LLM4THP as well as related peptide prediction frameworks, protein language model (PLM) embeddings combined with deep learning have shown promise.As demonstrated by models like PLMTHP and LLM4THP, as well as by related peptide prediction frameworks like PLMACPred for anticancer peptides and PepCNN for peptide–protein interactions, protein language model (PLM) embeddings combined with deep learning have demonstrated promise. Nevertheless, a lot of these approaches either use restricted feature combinations or concentrate on tasks related to THP prediction, which leaves space for more accurate and complex sequence pattern capture. We suggest a GAN-enhanced bidirectional semi-temporal CNN architecture that combines deep learning with PLM-based descriptors for THP prediction in order to close these gaps.As proven by models like PLMTHP and LLM4THP, and additionally by related peptide prediction frameworks like PLMACPred for anticancer peptides and PepCNN for peptide–protein interactions, protein language model (PLM) embeddings combined with deep learning have demonstrated promise. Nevertheless, a lot of these approaches either use restricted feature combinations or concentrate on tasks related to THP prediction, which leaves space for more accurate and complex sequence pattern capture. We suggest a GAN-enhanced bidirectional semi-temporal CNN architecture that combines deep learning with PLM-based descriptors for THP prediction in order to close these gaps.In conclusion, our work provides a cutting-edge computational tool for THP identification that connects biological understanding with computational prediction. The high accuracy and interpretability of the suggested model highlight the wider impact of AI-driven peptide discovery in precision oncology and support its potential as a basis for peptide-based drug delivery systems and cancer therapeutics development.
Supervised by Dr. Farman Ali
</summary>
<dc:date>2026-01-01T00:00:00Z</dc:date>
</entry>
<entry>
<title>Prediction of Epigenetic Protein using Deep Learning and Protein Large Language Models</title>
<link href="http://hdl.handle.net/123456789/21659" rel="alternate"/>
<author>
<name>Kaleem Ullah, 01-243241-010</name>
</author>
<id>http://hdl.handle.net/123456789/21659</id>
<updated>2026-08-31T09:46:20Z</updated>
<published>2026-01-01T00:00:00Z</published>
<summary type="text">Prediction of Epigenetic Protein using Deep Learning and Protein Large Language Models
Kaleem Ullah, 01-243241-010
Epigenetic proteins (EPs) are essential in regulating gene activity without altering the DNA sequence. Their role in numerous diseases, including cancer, neurodegenerative, and autoimmune disorders, has rendered them significant targets in drug discovery. It’s important to be able to accurately predict EPs because they play such a key role. Unfortunately, experimental methods for prediction are often costly and take a long time and required domain experts. In this study, we introduced a new deep learning framework that uses protein sequence data to make predictions about EPs. We obtained protein sequence data from Uniprot database and used two powerful pre-trained large language models, ProtBERT and UniREP, to turn protein sequences into rich numerical representations. Then, these embeddings are used as inputs for three different deep learning models: a Multi-Headed Ensemble Residual Convolutional Neural Network (MERCNN), a Bi-Directional Gated Recurrent Unit (Bi-GRU), and an Encoder-Decoder model. We systematically evaluated each model’s performance using both embedding models to determine the best combinations for predicting EP. Our model (MERCNN-EP) achieved an impressive 96.89% accuracy on the training dataset and 91.12% on the testing dataset, which was obtained from the UniProt database to evaluate the generalization ability of the proposed model. These results outperform previously published studies on epigenetic protein prediction, demonstrating the effectiveness and improved predictive capability of the proposed framework. Our results show that the MERCNN architecture works best when combined with ProtBERT embeddings. This research creates a strong and scalable framework for high-capacity EP screening, which could have a big impact on improving targeted therapeutic strategies in precision medicine.
Supervised by Dr. Farman Ali
</summary>
<dc:date>2026-01-01T00:00:00Z</dc:date>
</entry>
<entry>
<title>Exploring Evolutionary Algorithms for Optimal Features Selection to Detect Anomaly Based Intrusion In IoT</title>
<link href="http://hdl.handle.net/123456789/19737" rel="alternate"/>
<author>
<name>Ali Arshad, 01-243231-002</name>
</author>
<id>http://hdl.handle.net/123456789/19737</id>
<updated>2025-10-08T03:57:48Z</updated>
<published>2025-01-01T00:00:00Z</published>
<summary type="text">Exploring Evolutionary Algorithms for Optimal Features Selection to Detect Anomaly Based Intrusion In IoT
Ali Arshad, 01-243231-002
In light of the fast-evolving scenario of the IoT, network security constantly bears immense challenges due to the increasing number of cyber threats and vulnerabilities. Traditional intrusion detection systems use predefined signatures and rule-based approaches to detect malicious activities within a network. In contrast to the traditional signature-based intrusion detection system, which utilizes previously established attack patterns, an anomaly-based intrusion detection system typically utilizes machine learning, statistical models, and artificial intelligence to assess network traffic, system log entries, and user behavior. . This research proposes the anomaly-based intrusion detection system using Particle Swarm Optimization (PSO)-based feature selection and ensemble learning in enhancing detection accuracy for IoT networks. Performance of stacking, hard voting, soft voting, and autoencoder-based models is evaluated over the benchmark datasets NSL-KDD and KDDCup 99 in analyzing their effectiveness in detecting anomalous behaviors in IoT environments. From the results, it is clear that PSO-based feature selection is highly significant in anomaly detection. Anomaly detection gets better performance by reducing feature redundancy along with improving classification accuracy. Of all the models tested, Stacking performed the best, with an accuracy of 98.87% on NSL-KDD and 99.76% on KDDCup 99, proving to be the most effective method. Soft Voting and Hard Voting also did well on NSL-KDD, recording 98.38% and 98.13% accuracy respectively, which highlighted the strength of ensemble methods in identifying anomalies based on IoT. The Autoencoder model demonstrated unsupervised anomaly detection, with a slightly less accurate of 96.63% on NSL-KDD and 98.23% on KDDCup 99, due to its greater false positive rates. This suggests that models based on deep learning could make a significant performance improvement on anomaly classification by leveraging their explicit feature selection capabilities. Stacking is an ensemble learning algorithm that increases predictive performance by aggregating multiple base models in a superior way. It can be used optimally when there exists some labeled data in supervised learning cases. Auto-Encoders correspondingly are neural networks deployed mainly in unsupervised learning, and they serve very well in detecting anomalies without labeled data.
Supervised by Dr. Saba Mahmood
</summary>
<dc:date>2025-01-01T00:00:00Z</dc:date>
</entry>
</feed>
