
Revolutionising spatial multi-omics: HKBU's Computer Science disentangles complex biological data

Spatial multi-omics technologies have fundamentally transformed our understanding of complex biological systems. By capturing spatially resolved profiles from genomics, transcriptomics, proteomics, and metabolomics within their native tissue architecture, researchers can gain a comprehensive view of cellular states. However, combining these intricate datasets has historically presented a significant analytical challenge. Existing integration methods typically force data from different omics into a single, unified latent space, which inherently obscures the unique, modality-specific insights that each layer provides, severely limiting the full potential of multi-omics analyses. To overcome this critical barrier, Professor Zhang Lu, Associate Professor at the Department of Computer Science, led a team that included Professor Han Bo to develop the Spatial Multi-View (SpaMV) representation learning algorithm.
The dual-space architecture of SpaMV
“Unlike traditional methods that conflate distinct data types, SpaMV explicitly separates cross-omics shared information from omics-specific private information into entirely distinct latent spaces,” said Professor Zhang. The algorithm accomplishes this robust spatial multi-omics decomposition through three foundational components. First, it utilises a dual-encoder architecture relying on graph attention networks (GATs). Each omics layer employs both a shared encoder and a private encoder to extract spatially informed embeddings, alongside an omics-specific decoder to ensure accurate data reconstruction.
Second, SpaMV integrates auxiliary measurement models especially designed to minimise the mutual information between a private latent variable from one omics and the data originating from another. This mechanism actively prevents the leakage of shared biological information into private spaces. Finally, the algorithm enforces strict statistical independence between shared and private variables using a non-parametric independence test known as the Hilbert-Schmidt Independence Criterion, which ensures that the shared spatial representation is never contaminated by omics-specific anomalies.
Outperforming existing integration models
The efficacy of SpaMV was extensively benchmarked against eight state-of-the-art spatial integration methods, including SpatialGlue, CellCharter, SMOPCA, and COSMOS. Across three simulated datasets and five real-world datasets, spanning transcriptome-epigenome, transcriptome-metabolome, and transcriptome-proteome combinations, SpaMV consistently demonstrated a superior spatial domain clustering performance.
For example, when applied to clear cell renal cell carcinoma (ccRCC) spatial transcriptome-metabolome data, SpaMV achieved distinct and spatially coherent domain boundaries. The algorithm uncovered metabolome-specific spatial programmes, such as stress-associated lipid storage characterised by triglycerides, which were completely invisible when analysing transcriptomic data alone. Similarly, in mouse embryo datasets co-profiling spatial histone modifications, SpaMV successfully isolated anatomically distinct structures, such as the jaw, tongue, and nose regions, which competing algorithms either blended together or missed entirely.
Enhanced interpretability and biomarker discovery
“Beyond superior clustering, SpaMV provides easily interpretable topic modelling. It allows researchers to visualise both shared and private topics, making it abundantly clear which tissue segments are defined by a multi-omics consensus, and which are driven by a specific, isolated omics layer,” said Professor Zhang.
This capability is critical for accurate cell type annotation and avoiding misleading biological conclusions. In a spatially resolved mouse thymus dataset containing paired transcriptomics and proteomics, SpaMV identified a proteome-specific topic characterised by classic B cell markers like CD45R and CD19. Standard integration methods that rely solely on transcriptomic differential expression to define spatial domains often mislabel such regions, leading to false biomarker identification. Because SpaMV disentangles these signals, researchers can confidently utilise targeted proteomics information to annotate spatial domains, avoiding these common analytical pitfalls. By preserving the integrity of diverse data types, SpaMV provides an extremely accurate, systems-level view of complex tissue heterogeneity, and paves the way for deeper biological discoveries.
Full paper on Nature Communications: https://www.nature.com/articles/s41467-026-74718-1
Professor Zhang’s research profile: https://scholars.hkbu.edu.hk/en/persons/ERICLUZHANG
.png)
Professor Zhang Lu
Faculty of Science and Technology
Next News


