Exploring independent omics data fusion for precision oncology: cancer genes, biomarkers, and machine learning
| dc.contributor.advisor | Minghim, Rosane | |
| dc.contributor.advisorexternal | Leme, Adriana Franco Paes | |
| dc.contributor.author | Younis, Haseeb | en |
| dc.contributor.funder | Research Ireland | |
| dc.date.accessioned | 2026-01-21T09:26:18Z | |
| dc.date.available | 2026-01-21T09:26:18Z | |
| dc.date.issued | 2025 | |
| dc.date.submitted | 2025 | |
| dc.description.abstract | Cancer is a multifaceted and heterogeneous disease that continues to pose a major global health challenge. Despite technological advances, accurate classification of cancer types and subtypes, discovery of reliable molecular biomarkers remain difficult due to the biological complexity and diversity of patients and tumour types. With the emergence of high-throughput technologies, large scale omics datasets have become available, capturing distinct molecular layers of cancer, such as genomic mutations, gene expression, and protein interactions. While many studies attempt to integrate multiple omics, this work emphasizes the importance of understanding the strengths, limitations, and biological markers within each modality independently. Such focused analysis enables the design of data-specific algorithms with clearer interpretation, helping reveal complementary insights from each omics layer. This thesis presents novel computational approaches to the investigation of three independent omics data types for cancer classification and biomarker discovery: somatic mutation data (genomics), RNA Sequencing (RNA-seq) and microarray-based gene expression data (transcriptomics), and Mass Spectrometry (MS)-based protein co-fractionation data (proteomics). Each data type is investigated separately using tailored machine learning and deep learning algorithms, allowing precise modelling and interpretation of the biological and computational challenges unique to each dataset. With regards to genomics, a pipeline termed “CLARITY” was proposed, to desparsify somatic mutation data by summarising mutations into pathway scores, followed by clustering with autoencoders to assign cancer stage labels. Classical machine learning models, as well as newly designed domain-specific deep learning models, were employed to classify cancer stages. The results demonstrated high accuracy. Furthermore, survival analysis validated the biological relevance of the identified subtypes and potential driver genes. For transcriptomic analysis, deep learning architectures, specifically Convolutional Neural Network (CNN) and transformer models, were developed for cancer subtype classification with a focus on RNA-seq and microarray datasets from multiple cancer types. Synthetic Minority Over-sampling Technique (SMOTE) was used to address class imbalance, while Local Interpretable Model-agnostic Explanations (LIME) helped interpret model outputs and identify gene-level biomarkers. The models achieved competitive performance across various datasets and uncovered interpretable gene expression signatures with biological significance; this model was further validated through pathway enrichment analysis. The proteomic section of this research involved MS-based co-fractionation data collected from normal and cancerous oral keratinocyte cell lines. A ranking pipeline using feature selection was implemented to select robust protein subsets. Machine learning classifiers were then trained to differentiate between normal, less aggressive, and highly aggressive cancer lines, preserving the important proteins. Protein–Protein Interaction Network (PPIN)s were constructed to identify disrupted complexes specific to aggressive phenotypes, offering insight into molecular mechanisms underlying oral squamous cell carcinoma progression. Throughout this thesis, the challenges of high dimensionality, class imbalance, and data heterogeneity are addressed using domain-appropriate preprocessing, feature selection, and Explainable Artificial Intelligence (XAI) techniques. Rather than relying solely on performance metrics, a strong emphasis was placed on biological interpretability, making the findings valuable for both computational researchers and biomedical scientists. Collectively, this thesis demonstrates that each omics modality offers a distinct and valuable lens for understanding cancer biology. Genomic data provides information about mutation-driven mechanisms and tumour staging; transcriptomic data enables subtype classification and functional gene identification; proteomic data reveals disruptions in protein abundance and complex formation. An in-depth examination of these layers independently establishes a strong basis for future work aiming to integrate them. By providing validated models and feature sets, this thesis contributes to the development of interpretable, data-driven cancer classification frameworks and biomarker discovery pipelines. The methodologies and findings presented here lay the groundwork for more generalizable and biologically meaningful applications of machine learning in computational oncology, ultimately advancing efforts toward precision cancer diagnosis and therapy. | en |
| dc.description.status | Not peer reviewed | en |
| dc.description.version | Accepted Version | en |
| dc.format.mimetype | application/pdf | en |
| dc.identifier.citation | Younis, H. 2025. Exploring independent omics data fusion for precision oncology: cancer genes, biomarkers, and machine learning. PhD Thesis, University College Cork. | |
| dc.identifier.endpage | 178 | |
| dc.identifier.uri | https://hdl.handle.net/10468/18421 | |
| dc.language.iso | en | |
| dc.publisher | University College Cork | en |
| dc.relation.project | Research Ireland (Grant no. 18/CRT/6222) | |
| dc.rights | © 2025, Haseeb Younis. | |
| dc.rights.uri | https://creativecommons.org/licenses/by/4.0/ | |
| dc.subject | Cancer genes | |
| dc.title | Exploring independent omics data fusion for precision oncology: cancer genes, biomarkers, and machine learning | |
| dc.type | Doctoral thesis | en |
| dc.type.qualificationlevel | Doctoral | en |
| dc.type.qualificationname | PhD - Doctor of Philosophy | en |
Files
Original bundle
1 - 3 of 3
Loading...
- Name:
- YounisH_PhD2025.pdf
- Size:
- 27.41 MB
- Format:
- Adobe Portable Document Format
- Description:
- Full Text E-thesis
Loading...
- Name:
- YounisH_PhD2025.zip
- Size:
- 18.86 MB
- Format:
- http://www.iana.org/assignments/media-types/application/zip
- Description:
- Zip File
Loading...
- Name:
- Final-Thesis-Submission-Examination-Mr-Haseeb-Younis.pdf
- Size:
- 3.56 KB
- Format:
- Adobe Portable Document Format
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 5.2 KB
- Format:
- Item-specific license agreed upon to submission
- Description:
