❌

Normal view

scCluBench: Comprehensive Benchmarking of Clustering Algorithms for Single-Cell RNA Sequencing

arXiv:2512.02471v1 Announce Type: cross Abstract: Cell clustering is crucial for uncovering cellular heterogeneity in single-cell RNA sequencing (scRNA-seq) data by identifying cell types and marker genes. Despite its importance, benchmarks for scRNA-seq clustering methods remain fragmented, often lacking standardized protocols and failing to incorporate recent advances in artificial intelligence. To fill these gaps, we present scCluBench, a comprehensive benchmark of clustering algorithms for scRNA-seq data. First, scCluBench provides 36 scRNA-seq datasets collected from diverse public sources, covering multiple tissues, which are uniformly processed and standardized to ensure consistency for systematic evaluation and downstream analyses. To evaluate performance, we collect and reproduce a range of scRNA-seq clustering methods, including traditional, deep learning-based, graph-based, and biological foundation models. We comprehensively evaluate each method both quantitatively and qualitatively, using core performance metrics as well as visualization analyses. Furthermore, we construct representative downstream biological tasks, such as marker gene identification and cell type annotation, to further assess the practical utility. scCluBench then investigates the performance differences and applicability boundaries of various clustering models across diverse analytical tasks, systematically assessing their robustness and scalability in real-world scenarios. Overall, scCluBench offers a standardized and user-friendly benchmark for scRNA-seq clustering, with curated datasets, unified evaluation protocols, and transparent analyses, facilitating informed method selection and providing valuable insights into model generalizability and application scope.

Integrative Analysis of Multi-Omics Data for Biomarker Discovery

Annu Int Conf IEEE Eng Med Biol Soc. 2025 Jul;2025:1-7. doi: 10.1109/EMBC58623.2025.11254134.

ABSTRACT

The complexity of biological systems and the limitations of analyzing individual omics studies for biomarker discovery have raised the need for a holistic approach by multi-omics integration. By integrating data from multiple layers, researchers can gain insights into the entire system rather than just individual components. Also, integrative analysis can help identify molecular signatures that are more accurate in predicting disease onset, progression, and response to treatment, leading to better-targeted therapies and personalized medicine. In this paper, we explored statistical and deep learning methods for integrative analysis of metabolomics, lipidomics, peptidomics, proteomics, and glycoproteomics data acquired by LC-MS/MS analysis of serum samples from 20 hepatocellular carcinoma (HCC) cases and 20 patients with liver cirrhosis (CIRR). The goal is to identify a panel of multi-omics features that distinguish HCC cases from cirrhotic controls. A pathway analysis using these features identified biological pathways such as LXR/RXR Activation and Acute Response signaling as significantly enriched in our multi-omics datasets.

PMID:41336317 | PMC:PMC12694951 | DOI:10.1109/EMBC58623.2025.11254134

Multi-Omics Feature Selection to Identify Biomarkers for Hepatocellular Carcinoma

Metabolites. 2025 Aug 28;15(9):575. doi: 10.3390/metabo15090575.

ABSTRACT

INTRODUCTION: Hepatocellular carcinoma (HCC), the most prevalent form of liver cancer, ranks as the third leading cause of mortality globally. Patients diagnosed with HCC exhibit a dismal prognosis mostly due to the emergence of symptoms in the advanced stages of the disease. Moreover, conventional biomarkers demonstrate insufficient efficacy in the early detection of HCC, hence highlighting the need for the identification of novel and more effective biomarkers.

METHODS: In this paper, we investigate methods for integration of multi-omics data we generated by both untargeted and targeted mass spectrometric analysis of serum samples from HCC cases and patients with liver cirrhosis. Specifically, the performances of several feature selection methods are evaluated on their abilities to identify a panel of multi-omics features that distinguish HCC cases from cirrhotic controls.

RESULTS: The integrative analysis identified key molecules associated with liver including such as leucine and isoleucine as well as SERPINA1, which is involved in LXR/RXR Activation and Acute Response signaling. A new method that uses recursive feature selection in conjunction with a transformer-based deep learning model as an estimator led to more promising results compared to other deep learning methods that perform disease classification and feature selection sequentially.

CONCLUSIONS: The findings in this study reinforce the importance of adapting or extending deep learning models to support robust feature selection, especially for integration of multi-omics data with limited sample size to avoid the risk of overfitting and the need for evaluation of the multi-omics features discovered in this study via blood samples from a larger and independent cohort to identify robust biomarkers for HCC.

PMID:41002959 | PMC:PMC12471784 | DOI:10.3390/metabo15090575

❌