Multi-Modal Autoencoder Framework for Predicting Plasma Proteome from Genomic Data Using AlphaGenome
This study aims to investigate the integration of multi-modal genomic data to enhance the prediction of plasma protein abundance with AlphaGenome predictions using the UK Biobank Pharma Proteomics Project dataset, with a focus on cancer genetic variants and their proteomic effects. The primary objective is to develop a deep learning autoencoder model that combines genomic variant data from cancer GWAS with AlphaGenome derived regulatory features to predict how genetic variation influences protein expression levels in cancer contexts [1,2]. The model will learn compressed representations (embeddings) that capture how genetic variants influence protein expression through gene regulatory mechanisms, addressing the fundamental challenge of predicting molecular phenotypes from genotypic data in cancer biology. Assigning function to genetic variants as expression quantitative trait loci is an expanding and useful approach, but focuses exclusively on mRNA rather than protein levels, and many variants remain without annotation [3]. Recent work has demonstrated that trans-associations of cancer related variants with distal molecular targets can identify convergent regulatory effects across related cancers, with studies identifying shared proteomic effects across cancer types [4], highlighting the importance of understanding trans-regulatory mechanisms in protein abundance prediction for cancer research. By incorporating state-of-the-art AI predictions of intermediate regulatory processes through AlphaGenome [5], this work aims to bridge the gap between cancer genotype and proteotype more effectively than traditional approaches.
The student will design and implement a neural network architecture that processes multi-modal genomic data from cancer GWAS, combining genotype information from genome significant cancer associated pQTL variants with AlphaGenome's functional predictions. AlphaGenome takes as input DNA sequence and predicts thousands of functional genomic tracks [6]. The focus on cancer associated variants will enable the model to learn regulatory patterns specific to oncogenic processes and tumor biology. A deeply stacked denoising autoencoder approach has been successfully applied to protein structure reconstruction [6], demonstrating the viability of deep autoencoder architectures for biological prediction tasks. Performance evaluation will compare architectural variants with benchmarking against traditional linear regression approaches used in conventional cancer pQTL analyses [7,8].
The final phase focuses on extracting biological insights relevant to cancer biology through systematic interpretation and validation of predictions. The student will implement attention mechanisms or gradient-based feature attribution methods to quantify which regulatory features contribute most to accurate protein abundance predictions in cancer contexts. Protein quantitative trait loci (pQTLs) identify novel relationships between inter-individual protein levels and genetic variants [3], and understanding which regulatory mechanisms mediate these relationships is crucial for identifying potential therapeutic targets and biomarkers in cancer. Latent space analysis using dimensionality reduction will reveal whether the model learns biologically meaningful clusters where cancer-associated variants with similar regulatory mechanisms group together, potentially identifying shared regulatory programs across cancer types [9].
Validation will include stratified performance analysis comparing cis-pQTLs versus trans-pQTLs in cancer-related genes, testing whether incorporating regulatory information particularly improves predictions for non-coding variants in cancer-relevant regulatory regions, and conducting case studies on well-characterized oncogenic regulatory variants to confirm if the model captures known biological mechanisms in cancer pathways. Expression quantitative trait locus (eQTLs) and other molecular QTL studies have been valuable resources in identifying candidate causal genes from GWAS loci through statistical colocalization methods [10], and this work extends that framework to the cancer proteome level with regulatory context.
Requisites
Implementation will use Python with PyTorch or TensorFlow, requiring development of efficient data preprocessing pipelines for large-scale cancer genomic data and feature engineering from AlphaGenome's outputs.
Cooperation with company or external entity: Potential collaboration with NCI’s (National Cancer Institute) researchers.
[1] Sun BB, et al. Plasma proteomic associations with genetics and health in the UK Biobank. Nature 622:329-338 (2023).
[2] Sun BB, et al. Genomic atlas of the human plasma proteome. Nature 558:73-79 (2018).
[3] Lourdusamy A, et al. Protein Quantitative Trait Loci Identify Novel Candidates Modulating Cellular Response to Chemotherapy. PLoS Genetics 10(4):e1004192 (2014). PMC3974641.
[4] Mukherjee D, et al. TASTE identifies shared proteomic effects on multiple related cancers. medRxiv doi:10.64898/2025.12.19.25342717 (2025).
[5] Avsec Ž, et al. Advancing regulatory variant effect prediction with AlphaGenome. Nature 649:1206-1218 (2026).
[6] Li H, et al. A Template-Based Protein Structure Reconstruction Method Using Deep Autoencoder Learning. J Proteomics Bioinform 9(12):306-313 (2016). PMC29081613.
[7] Suhre K, et al. Connecting genetic risk to disease end points through the human blood plasma proteome. Nat Commun 8:14357 (2017).
[8] Melzer D, et al. A Genome-Wide Association Study Identifies Protein Quantitative Trait Loci (pQTLs). PLoS Genet 4(5):e1000072 (2008).
[9] Dutta D, et al. Aggregative trans-eQTL analysis detects trait-specific target gene sets in whole blood. Nat Commun 13:4323 (2022).
[10] Zhang Z, et al. ezQTL: Interactive Visualization and Colocalization of Quantitative Trait Loci and GWAS. NCI Division of Cancer Epidemiology and Genetics (2022). https://dceg.cancer.gov/tools/analysis/ez-qtl