cellart-turns-blurry-spatial-gene-maps-into-crisp-single-cell-atlases
CellART turns blurry spatial gene maps into crisp single-cell atlases

CellART turns blurry spatial gene maps into crisp single-cell atlases

Spatial transcriptomics has been one of the most celebrated breakthroughs in modern biology, a method so transformative that Nature Methods named it Method of the Year in 2021 and Nature listed it among the technologies to watch in 2024. The promise is seductive: instead of grinding tissue into a soup and losing all sense of geography, researchers can measure gene activity while preserving the precise location of every molecule inside a slice of tumor, brain or skin. Yet a stubborn technical gap has persisted between what the newest instruments record and what biologists actually need. Now a team at The Hong Kong University of Science and Technology, working with collaborators at Sun Yat-sen University, has unveiled a computational framework called CellART that promises to close that gap, converting high-resolution spatial data into clean, single-cell information across a remarkable range of platforms.

The problem CellART attacks is fundamental to how modern spatial platforms work. Instruments such as 10x Genomics’ Xenium and Visium HD, Vizgen’s MERFISH and BGI’s Stereo-seq can now resolve gene expression at subcellular scale, but they pay a price for that resolution. Some capture only sparse transcript counts per spot, while others measure a limited panel of genes. In imaging-based assays, transcripts are detected as individual dots scattered across a tissue image, and deciding which dots belong to which cell is far from trivial. On array-based platforms like Visium HD, the data arrive in tiny bins that frequently straddle cell boundaries, so a single bin may contain RNA from two or three neighboring cells. Without a reliable way to draw cell boundaries and assign transcripts, the single-cell picture that biologists crave remains frustratingly blurred.

CellART, described in a Brief Communication published in Nature Computational Science on 5 October 2026, tackles both challenges at once. The framework performs cell segmentation and cell-type annotation simultaneously, fusing three complementary data streams: the staining images that accompany spatial assays, the spatial transcriptomics data themselves, and single-cell RNA sequencing references that catalog the expression signatures of known cell types. The key insight is that these modalities carry mutually reinforcing information. Histology images reveal where nuclei and membranes lie; transcript distributions show where gene products accumulate; and reference atlases constrain what combinations of cell types are biologically plausible in a given tissue. By integrating deep learning with probabilistic modeling, CellART lets each data source correct the weaknesses of the others.

Technically, the segmentation module builds on ideas from computer vision, drawing on architectures such as feature pyramid networks and residual networks that have proven their worth in object detection and image recognition. These networks learn to delineate cell boundaries from the staining images and the spatial patterns of transcript density. On top of that segmentation layer sits a probabilistic model in the spirit of deep generative approaches to single-cell transcriptomics, which assigns each segmented cell a cell-type label by comparing its expression profile with the scRNA-seq reference. This joint design means that segmentation errors and annotation errors can be caught and corrected in a single pass rather than compounding across separate pipeline stages, a common failure mode when researchers chain together independent tools.

The breadth of validation is one of the study’s most striking features. The team applied CellART to datasets spanning at least four distinct spatial platforms: Xenium, Visium HD, MERFISH and Stereo-seq, across tissues including human lung cancer, human breast cancer, human colorectal cancer, human skin, and mouse brain. On the Xenium 2.0 human lung dataset, where 10x Genomics provides a ground-truth segmentation derived from mixed-fluorescence imaging, CellART’s boundaries were quantified against established methods including Baysor, Cellpose and ProSeg using F1-scores and Jaccard indices. On Visium HD mouse brain data, where the recommended 16-micrometer bins do not align with cells in the histology image, CellART and Bin2Cell were compared by transcript coverage rate and by the correlation between cytoplasmic and nuclear expression profiles of individual cells, covering nearly sixty thousand cells in one benchmark alone.

Accuracy in cell typing proved equally important. In a human lung section profiled with both Xenium and post-Xenium Visium HD, CellART recovered consistent cell-type labels across the two platforms, suggesting that its annotations reflect genuine biology rather than platform-specific artifacts. On mouse brain data from all four platforms, the framework reconstructed the layered architecture of the cortex and reproduced the hippocampal organization documented in the Allen Mouse Brain Atlas, with the extracted cells correlating strongly with the single-cell reference. The team also measured computational efficiency by subsampling tissues of increasing size, and CellART remained tractable where some competing approaches struggled; one alternative method, TopACT, did not complete at full tissue size, a point the authors extrapolated rather than measured.

The biological payoff is vividly illustrated in cancer tissue. In the Xenium breast cancer dataset, CellART resolved ten cell types with high consistency between replicates, and its sharper boundaries prevented a classic pitfall: transcripts from neighboring cells bleeding into incorrectly drawn boundaries. Where a standard pipeline combining 10x segmentation with scANVI misassigned cells as tumor because ERBB2 transcripts from adjacent malignant cells drifted inside their borders, CellART produced crisper ERBB2 spatial patterns with less transcript diffusion. In breast cancer, where ERBB2 defines the HER2-positive subtype and guides therapy, that distinction between a tumor cell and its neighbor is not a cosmetic detail; it can change how a tumor’s molecular profile is read.

Perhaps the most compelling demonstration came from Visium HD colorectal cancer data, where CellART’s single-cell resolution outperformed spot-level deconvolution by RCTD, which classified the majority of more than 429,000 spots as doublets containing mixed cell populations. Working at true single-cell scale, CellART separated macrophage subpopulations that the coarser analysis blurred together. It distinguished SELENOP-positive macrophages scattered among tumor cells from SPP1-positive macrophages clustered at specific regions, one of which was recovered only by CellART. Defining tumor-associated macrophages as macrophages lying within fifty micrometers of a tumor cell, the team found these cells concentrating along tumor boundaries and forming a transcriptionally distinct cluster between macrophages and tumor cells, with upregulated immune-related and tumor-promoting genes, particularly MMP12.

That last finding hints at the kind of discovery CellART makes routine. Ligand-receptor analysis revealed strong interactions between tumor-associated macrophages and tumor cells, including the MMP12-PLAUR, SPP1-ITGAV and ITGB5, and APOE-SORL1 pairs, with their spatial expression co-localized at tumor margins. The urokinase plasminogen activator system, to which PLAUR belongs, has long been implicated in cancer metastasis and prognosis, and integrins are established drivers of invasion. Seeing these interaction axes light up precisely where macrophages and malignant cells touch provides a spatially grounded view of the tumor microenvironment that neither dissociated single-cell sequencing nor coarse spatial deconvolution can deliver. For immunotherapy research, where the dialogue between immune cells and tumor cells is everything, tools of this kind could reshape how tumor sections are interpreted.

Accessibility may determine whether CellART achieves the adoption its designers hope for. The software is written in Python, released under the permissive MIT licence on GitHub, and archived on Zenodo, with simulation and evaluation scripts, documentation and worked examples included. Its outputs are compatible with widely used community tools, including the SpatialData framework for spatial omics, easing integration into existing analysis workflows. All the spatial and single-cell reference datasets analyzed in the study are publicly available from 10x Genomics, Vizgen, BGI, the Allen Institute, the CZI CELLxGENE portal and the Allen Brain Atlas, and the processed inputs and simulated benchmark data are deposited for reuse. In a field where proprietary pipelines and incompatible file formats often fragment progress, a unified, open, platform-agnostic tool that turns raw high-resolution spatial data into trustworthy single-cell atlases could become as indispensable as the sequencing machines that generate the data in the first place.

Subject of Research: A unified computational framework for single-cell segmentation and annotation in high-resolution spatial transcriptomics

Article Title: CellART: a unified framework for extracting single-cell information from high-resolution spatial transcriptomics

Article References: CellART: a unified framework for extracting single-cell information from high-resolution spatial transcriptomics. (n.d.). https://doi.org/10.1038/s43588-026-01054-1

Image Credits: AI Generated

DOI: 10.1038/s43588-026-01054-1

Keywords: spatial transcriptomics, CellART, single-cell analysis, cell segmentation, deep learning, probabilistic modeling, tumor microenvironment, macrophages, Xenium, Visium HD, MERFISH, computational biology