Understanding how cancers interact with the surrounding immune system is central to developing better treatments. The tumor microenvironment (TME) -- the complex mixture of immune cells, supporting cells, and signaling molecules surrounding a tumor -- plays a decisive role in whether a cancer grows, responds to treatment, or evades the immune system entirely.
Traditional methods for studying the TME, such as flow cytometry or bulk RNA sequencing, average the gene activity of thousands of cells together. While useful, this averaging obscures critical differences between rare cell subpopulations that may be essential for understanding treatment resistance or immune evasion.
Single-cell RNA sequencing (scRNA-seq) measures gene expression in each individual cell separately, generating a molecular portrait of every cell in a sample. This has transformed our ability to identify new cell types, track how cells change state during disease progression, and understand how tumor cells communicate with immune cells.
Over 1,400 bioinformatics tools now exist for analyzing single-cell data, and this review provides a practical guide to the computational methods used to extract meaningful biological information from these complex datasets -- and how that information is beginning to shape cancer treatment decisions.
Two main approaches exist for single-cell RNA sequencing. Plate-based methods (such as SMART-seq2) use cell sorting to place individual cells into wells, allowing full-length gene sequencing. This is ideal for detecting rare cell types but can only process hundreds of cells per experiment.
Droplet-based methods (such as 10x Chromium) encapsulate thousands of individual cells in tiny fluid droplets, each tagged with a unique barcode. This enables much higher throughput -- tens of thousands of cells per run -- making large-scale tumor studies practical, though at the cost of only sequencing partial gene transcripts.
A newer category of spatial transcriptomics technologies preserves the physical location of cells within a tissue while measuring their gene expression. Methods like Visium, MERFISH, and Slide-seq allow researchers to map not just what genes are active in each cell, but where in the tumor those cells are located -- revealing how different immune populations are distributed across tumor regions.
The most advanced platforms now simultaneously measure both RNA and protein markers at single-cell resolution while preserving spatial location (CosMx, Xenium, PhenoCycler). These multi-modal approaches provide the most comprehensive picture of the tumor immune microenvironment, though they come with significant cost and computational complexity.
Raw single-cell sequencing data requires extensive processing before biological insights can be extracted. Quality control is the first step: tools like SoupX and CellBender remove contamination from RNA that leaked out of cells (ambient RNA), while scDblFinder eliminates doublets -- cases where two cells were accidentally captured together and appear as one.
Normalization corrects for the fact that some cells were sequenced more deeply than others. The regression-based sctransform approach has become a standard choice because it treats sequencing depth as a statistical covariate, reducing technical noise while preserving true biological variation. Batch correction tools like Harmony and scVI then align data from multiple samples or experiments so they can be analyzed together.
Dimensionality reduction compresses the thousands of gene measurements per cell into two or three dimensions for visualization and clustering. Linear methods like PCA are fast but miss non-linear relationships. Non-linear methods like UMAP (Uniform Manifold Approximation and Projection) and t-SNE better preserve the complex biological structure of single-cell data and have become the standard for visualizing cell populations.
Cell annotation -- determining what cell type each cluster represents -- is often the most subjective step. Automated tools like CellTypist and Azimuth use reference datasets to classify cells automatically, but expert review remains essential, particularly for novel or unusual cell states not well-represented in existing reference databases.
One of the most powerful applications of single-cell data is mapping cell-cell communication (CCC) -- determining which cells send signals to which other cells and through which molecular pathways. Most CCC analysis infers interactions from matching ligand-receptor pairs: if cell type A expresses a signaling molecule (ligand) and cell type B expresses its matching receptor, an interaction is inferred.
Tools like CellChat and CellPhoneDB use statistical permutation tests to determine which ligand-receptor interactions are genuinely elevated above background. CellChat goes further by incorporating the effects of co-receptors and signaling mediators, and applying network analysis to identify which cell types are the dominant signal senders versus receivers in a tumor.
NicheNet and similar tools extend this analysis downstream -- not just asking which cells interact, but predicting which genes within a receiving cell are activated or suppressed by the incoming signals. This links surface-level interactions to intracellular gene regulatory changes, providing mechanistic insight into how immune cells are being influenced by their tumor environment.
A key limitation of current CCC tools is that they predict interactions based on RNA measurements, but actual protein binding occurs at the cell surface. Post-translational protein modifications (such as glycosylation) can dramatically alter receptor function without changing RNA levels. Future approaches combining RNA data with direct protein measurements will be needed for more accurate communication mapping.
Several studies highlighted in this review demonstrate the emerging clinical value of single-cell approaches for predicting how patients will respond to specific cancer treatments. In pancreatic cancer, a novel macrophage subtype called meCAF (metabolic cancer-associated fibroblast) was identified by single-cell analysis and found to be paradoxically associated with both worse prognosis and better response to anti-PD-1 immunotherapy -- a nuance that bulk sequencing would have missed.
In kidney cancer treated with immunotherapy, single-cell studies found that patients with high levels of CD8+ tissue-resident T cells before treatment showed better responses to immune checkpoint inhibitors. A specific macrophage subtype (ISGhi TAM) was also identified and validated across multiple independent patient cohorts as a predictor of improved outcomes following targeted therapy with sunitinib.
For lung cancer, 49 distinct immune cell populations were identified by single-cell analysis from 35 tumor samples. The relative abundance and interactions among activated T cells, antibody-producing plasma cells, and specific macrophages were combined into a single LCAM (lung cancer activation module) score. Patients with high LCAM scores showed significantly better responses to anti-PD-1 therapy.
These examples illustrate a common trajectory: single-cell biology identifies complex immune signatures, which are then validated in larger cohorts and simplified into clinically applicable scoring systems that can be measured using more accessible technologies like immunohistochemistry or targeted gene panels.
A persistent technical challenge is dropout -- the prevalence of zero gene expression values in single-cell data, caused by inefficient RNA capture during the sequencing process. Statistical and deep learning-based imputation tools help address this, but they can introduce biases if the underlying data patterns do not match their assumptions. Larger datasets and improved experimental capture efficiency are the most reliable solutions.
Integrating data from multiple studies or hospitals remains difficult due to batch effects -- systematic differences in measurements introduced by different laboratories, equipment, and protocols. Harmonization tools like Harmony and Scanorama help, but standardized experimental protocols and shared reference databases (like the Human Cell Atlas) are ultimately needed for robust cross-study comparisons.
Current scRNA-seq also provides only a snapshot of cell states at one moment in time, making it difficult to study dynamic processes. Combining sequencing with experimental perturbations (such as treating cells with drugs and measuring the response), temporal sampling, or trajectory inference algorithms helps reconstruct the dynamic evolution of cell states over time.
Looking ahead, the most promising direction is bench-to-bedside translation: the complex multi-dimensional cell atlas information from single-cell studies is being distilled into simple, clinically measurable panels of protein markers or gene scores that can be applied to routine tissue biopsies to guide treatment decisions in everyday oncology practice.