What is it about?
Cancer studies often search for important genes by looking for genes whose activity varies strongly across patients. But a gene does not need to show large variation to play an important structural role in a biological system. In this study, we analyzed gene-expression data from 526 breast cancer samples covering 17,800 genes. Instead of ranking genes only by how much their expression varies, we asked a different question: how much does a small change in one gene alter the overall structure of the data? We developed a manifold-based approach that maps thousands of gene measurements into a structured representation and then measures how sensitive that structure is to changes in individual genes. This revealed genes such as ASCL2 and NAPRT1 that ranked very low by conventional variance measures but showed strong structural influence. These genes also produced substantially better separation of patient survival groups than a conventional PCA-based approach.
Featured Image
Photo by Sasun Bughdaryan on Unsplash
Why is it important?
Modern cancer datasets contain tens of thousands of genes, so researchers must decide which signals deserve further attention. Common filtering methods can remove genes simply because their expression does not vary enough across patients. Our results show that some of these apparently low-variance genes may still have unusually strong influence on the global structure of the data. This provides a complementary way to prioritize genes: not only by how much they change, but by how strongly a small perturbation propagates through the broader molecular representation. In the breast cancer dataset studied here, this approach identified hidden structural signals that conventional variance-based ranking missed and produced stronger survival stratification. The framework may therefore help researchers generate new hypotheses and identify candidate genes for further biological and clinical investigation, particularly in high-dimensional genomic datasets where important structural relationships can be obscured by conventional feature-selection methods.
Read the Original
This page is a summary of: Geometric Prognostic Singularities and Structural Leverage
Drivers: A Manifold-Based Framework (M-CIM) for Cancer Gene Prioritization, January 2026, Springer Science + Business Media,
DOI: 10.21203/rs.3.rs-8537800/v1.
You can read the full text:
Resources
Beyond Variance: Geometric Leverage for AI-Driven Gene Discovery in Cancer Transcriptomes
Open-access version of the study introducing a manifold-based geometric leverage framework for cancer gene discovery and transcriptomic analysis. The work analyzes TCGA Breast Invasive Carcinoma (TCGA-BRCA) gene-expression data and measures how perturbations to individual genes alter the global geometry of the patient-data manifold. This provides a structural sensitivity measure that complements conventional variance-based gene ranking and dimensionality-reduction approaches. The study identifies genes whose influence on the global transcriptomic structure may be substantial despite relatively low conventional variance rankings, including ASCL2 and NAPRT1. Downstream analyses examine whether these structurally influential genes provide useful information for patient stratification and survival separation. The preprint contains the complete methodology, mathematical formulation, computational experiments, gene-ranking results, validation analyses, and discussion of potential applications to high-dimensional genomics and AI-assisted gene discovery.
BRCA Multi-Omics Dataset and Analysis Pipeline
This related computational resource provides an analysis-ready breast cancer multi-omics dataset derived from publicly released TCGA GDAC data. It integrates median-normalized mRNA expression, mature miRNA expression (RPM), and Tier 1 clinical metadata within a common sample-aligned framework. The repository also includes executable analysis pipelines for sample harmonization, mRNA–miRNA feature fusion, dimensionality reduction using PCA, clustering, and exploratory multi-omics analysis. It is designed for methodological prototyping, machine-learning and deep-learning experiments, model benchmarking, and bioinformatics research. While this resource represents a separate computational workflow from the geometric leverage analysis described in the associated publication, both investigate high-dimensional structure in breast cancer molecular data and provide complementary approaches for studying feature relevance, patient structure, and computational cancer genomics.
Contributors
The following have contributed to this page







