What is it about?
A single cell can be described in two very different ways: by the genes it is switching on, and by the proteins sitting on its surface. New laboratory methods can measure both in the same cell at the same time, but the two measurements are hard to combine. There are tens of thousands of genes and only a hundred or so proteins, and the two kinds of data are noisy in different ways. Most existing tools therefore ask the user to decide how much weight each measurement should carry — a choice that changes the answer. We asked whether a simple linear algebra approach can combine the two without that choice. Applied to immune T cells from mice, it grouped cells sensibly and also pointed to the genes behind those groups, in one step and with far less computer time and memory than a comparable method. It did not beat protein-based methods at identifying cell types, and we say so: this is a lightweight, transparent alternative, not a replacement.
Featured Image
Photo by National Cancer Institute on Unsplash
Why is it important?
When a method requires the user to decide how much weight to give each type of measurement, that decision quietly shapes the result. Two researchers analysing the same cells can reach different conclusions simply because they set the balance differently, and there is usually no principled way to choose. Removing the choice removes one source of arbitrariness — and makes results easier to reproduce. The practical side matters too. Our approach ran on a standard computer using about a gigabyte of memory, several times faster than a comparable method, and needed no discarding of low-quality cells and no pre-filtering of genes beforehand. Because it is linear, it is also reversible: the genes contributing to the result can be traced afterwards, which is not possible with most deep-learning approaches. Equally important is what the method does not do. It did not outperform protein-based approaches at identifying cell types. Knowing where a method is useful, and where it is not, is more valuable to other researchers than a claim of general superiority.
Perspectives
I have spent many years applying tensor decomposition to problems where two very different kinds of measurement have to be looked at together — gene expression and methylation, messenger RNA and microRNA, and others. Each time, the same question comes up: how much should each measurement count? I have never found a satisfying answer, so I have preferred formulations that do not ask it. When single-cell methods began measuring genes and surface proteins in the same cell, it seemed natural to try the same construction. What I did not expect was how much of the work would end up being about honesty rather than performance. The method is fast, simple, and interpretable, but it does not win on every measure, and one dataset defeated it entirely. Reporting that felt more useful than quietly leaving it out. I hope readers take from this less a new tool to adopt than a reminder that a method's limits are part of its description.
Professor Y-h. Taguchi
Chuo Daigaku
Read the Original
This page is a summary of: Interpretable integration of CITE-seq RNA and ADT profiles without explicit modality-weight tuning via tensor decomposition-based unsupervised feature extraction, Scientific Reports, August 2026, Springer Science + Business Media,
DOI: 10.1038/s41598-026-64531-7.
You can read the full text:
Contributors
The following have contributed to this page







