preprint · bioRxiv (Cold Spring Harbor Laboratory) · 2018

Sparse network-based regularization for the analysis of patientomics high-dimensional survival data

André Veríssimo, Eunice Carrasquinha, Marta B. Lopes, Arlindo L. Oliveira, Marie‐France Sagot, Susana Vinga · 15 citations

View original publication →

See where this sits in the topic map →

Summary AI-generated

TL;DR
Modern sequencing technologies generate vast amounts of molecular data that make it difficult to build accurate and interpretable models for cancer survival analysis.
Problem
High-dimensional molecular data hampers the creation of survival models that are both accurate and easy to interpret.
Method
This work incorporates graph centrality measures into classical sparse survival models, such as the elastic net, using network information derived from external knowledge and the data itself.
Results
Introducing feature degree information into survival models consistently improves predictive performance on breast invasive carcinoma (BRCA) transcriptomic data from TCGA while enhancing model interpretability.
Contributions
Not specified in the abstract.
Limitations
Not specified in the abstract.
Takeaways
These approaches are implemented in the glmSparseNet R package, a flexible tool for exploring sparse network-based regularizers in generalized linear models for omics data analysis.
Applications
Preliminary clinical validation is supported through the Cancer Hallmarks Analytics Tool API and the String database.
Topics
Oncological survival analysis, network-based regularization, high-dimensional omics data, sparse generalized linear models
For industry
Healthcare and biomedical research
Why it matters
Advances computational methods for analyzing complex molecular data to support cancer research and survival modeling.

Abstract

Abstract Data availability by modern sequencing technologies represents a major challenge in oncological survival analysis, as the increasing amount of molecular data hampers the generation of models that are both accurate and interpretable. To tackle this problem, this work evaluates the introduction of graph centrality measures in classical sparse survival models such as the elastic net. We explore the use of network information as part of the regularization applied to the inverse problem, obtained both by external knowledge on the features evaluated and the data themselves. A sparse solution is obtained either promoting features that are isolated from the network or, alternatively, hubs, i.e. , features that are highly connected within the network. We show that introducing the degree information of the features when inferring survival models consistently improves the model predictive performance in breast invasive carcinoma (BRCA) transcriptomic TCGA data while enhancing model interpretability. Preliminary clinical validation is performed using the Cancer Hallmarks Analytics Tool API and the String database. These case studies are included in the recently released glmSparseNet R package 1 , a flexible tool to explore the potential of sparse network-based regularizers in generalized linear models for the analysis of omics data.

References within the group

← All publications