Sparse network-based regularization for the analysis of patientomics high-dimensional survival data
See where this sits in the topic map →Summary AI-generated
- TL;DR
- Modern sequencing technologies generate vast amounts of molecular data that make it difficult to build accurate and interpretable models for cancer survival analysis.
- Problem
- High-dimensional molecular data hampers the creation of survival models that are both accurate and easy to interpret.
- Method
- This work incorporates graph centrality measures into classical sparse survival models, such as the elastic net, using network information derived from external knowledge and the data itself.
- Results
- Introducing feature degree information into survival models consistently improves predictive performance on breast invasive carcinoma (BRCA) transcriptomic data from TCGA while enhancing model interpretability.
- Contributions
- Not specified in the abstract.
- Limitations
- Not specified in the abstract.
- Takeaways
- These approaches are implemented in the glmSparseNet R package, a flexible tool for exploring sparse network-based regularizers in generalized linear models for omics data analysis.
- Applications
- Preliminary clinical validation is supported through the Cancer Hallmarks Analytics Tool API and the String database.
- Topics
- Oncological survival analysis, network-based regularization, high-dimensional omics data, sparse generalized linear models
- For industry
- Healthcare and biomedical research
- Why it matters
- Advances computational methods for analyzing complex molecular data to support cancer research and survival modeling.
Abstract
Abstract Data availability by modern sequencing technologies represents a major challenge in oncological survival analysis, as the increasing amount of molecular data hampers the generation of models that are both accurate and interpretable. To tackle this problem, this work evaluates the introduction of graph centrality measures in classical sparse survival models such as the elastic net. We explore the use of network information as part of the regularization applied to the inverse problem, obtained both by external knowledge on the features evaluated and the data themselves. A sparse solution is obtained either promoting features that are isolated from the network or, alternatively, hubs, i.e. , features that are highly connected within the network. We show that introducing the degree information of the features when inferring survival models consistently improves the model predictive performance in breast invasive carcinoma (BRCA) transcriptomic TCGA data while enhancing model interpretability. Preliminary clinical validation is performed using the Cancer Hallmarks Analytics Tool API and the String database. These case studies are included in the recently released glmSparseNet R package 1 , a flexible tool to explore the potential of sparse network-based regularizers in generalized linear models for the analysis of omics data.