What is it about?

Cancer studies often search for important genes by looking for genes whose activity varies strongly across patients. But a gene does not need to show large variation to play an important structural role in a biological system. In this study, we analyzed gene-expression data from 526 breast cancer samples covering 17,800 genes. Instead of ranking genes only by how much their expression varies, we asked a different question: how much does a small change in one gene alter the overall structure of the data? We developed a manifold-based approach that maps thousands of gene measurements into a structured representation and then measures how sensitive that structure is to changes in individual genes. This revealed genes such as ASCL2 and NAPRT1 that ranked very low by conventional variance measures but showed strong structural influence. These genes also produced substantially better separation of patient survival groups than a conventional PCA-based approach.

Featured Image

Why is it important?

Modern cancer datasets contain tens of thousands of genes, so researchers must decide which signals deserve further attention. Common filtering methods can remove genes simply because their expression does not vary enough across patients. Our results show that some of these apparently low-variance genes may still have unusually strong influence on the global structure of the data. This provides a complementary way to prioritize genes: not only by how much they change, but by how strongly a small perturbation propagates through the broader molecular representation. In the breast cancer dataset studied here, this approach identified hidden structural signals that conventional variance-based ranking missed and produced stronger survival stratification. The framework may therefore help researchers generate new hypotheses and identify candidate genes for further biological and clinical investigation, particularly in high-dimensional genomic datasets where important structural relationships can be obscured by conventional feature-selection methods.

Read the Original

This page is a summary of: Geometric Prognostic Singularities and Structural Leverage Drivers: A Manifold-Based Framework (M-CIM) for Cancer Gene Prioritization, January 2026, Springer Science + Business Media,
DOI: 10.21203/rs.3.rs-8537800/v1.
You can read the full text:

Read

Resources

Contributors

The following have contributed to this page