❌

Normal view

Toward the best generalizable performance of machine learning in modeling omic and clinical data

17 October 2025 at 18:00

Lab Invest. 2025 Oct 15:104253. doi: 10.1016/j.labinv.2025.104253. Online ahead of print.

ABSTRACT

There are often performance differences between intra-dataset and cross-dataset tests in machine learning (ML) modeling. However, reducing these differences may reduce ML performances. It is thus a challenging dilemma for developing models that excel in intra-dataset testing and are generalizable to cross-dataset testing. Therefore, we aimed to understand and improve performance and generalizability of ML in intra-dataset and cross-dataset testing. We evaluated 4,200 ML models of classifying lung adenocarcinoma (LUAD) deaths using the The Cancer Genome Atlas (TCGA, n=286) and Oncogenomic-Singapore (OncoSG, n=167) datasets, and 1,680 models of classifying glioblastoma deaths using TCGA (n=151) and Clinical Proteomic Tumor Analysis Consortium (CPTAC, n=97) datasets. After examining performance distributions of these ML models, we applied a dual analytical framework, including statistical analyses and SHapley Additive exPlanations-based meta-analysis, to quantify factors' importance and trace model success back to design principles. We also developed a framework to identify the best generalizable model. Strikingly, Jarque-Bera test revealed significant deviations of model performances from normality in both cancer types and testing contexts. Simple linear models with sparse feature sets consistently dominated in LUAD experiments, whereas non-linear models dominated in glioblastoma ones, suggesting that the best modeling strategy appears cancer-type/disease dependent. Importantly, both robust Analysis of Variance (ANOVA) and Kruskal-Wallis tests consistently identified differentially expressed genes as one of the most influential factors in both cancer types. The proposed multi-criteria framework successfully identified the model that achieved both the best cross-dataset performance and similar intra-dataset performance. In summary, ML performance distributions significantly deviated from normality, which motivates using both robust parametric and non-parametric statistical tests. We quantified and provided possible exploitability on the factors associated with cross-dataset performances and generalizability of ML models in two cancer types. A multi-criteria framework was developed and validated to identify the models that are accurate and consistently robust cross datasets.

PMID:41106592 | DOI:10.1016/j.labinv.2025.104253

Integrated multi-omics analysis and experimental investigation of mitochondrial dynamics-related genes: molecular subtypes, immune landscape, and prognostic implications in lung adenocarcinoma

Front Immunol. 2025 May 29;16:1585505. doi: 10.3389/fimmu.2025.1585505. eCollection 2025.

ABSTRACT

BACKGROUND: Lung adenocarcinoma (LUAD) is a common and aggressive subtype of lung cancer associated with poor clinical outcomes. The role of mitochondrial dynamics (MD)-related genes in tumor progression and immune regulation remains poorly understood.

METHODS: Data from public databases were integrated, and subtypes were classified based on 23 MD-related genes. A five-gene prognostic model was constructed. Associations between the model and immune infiltration, tumor mutational burden (TMB), tumor stemness, and drug sensitivity were analyzed. The function of the key gene MTCH2 was validated through in vitro experiments.

RESULTS: Two distinct MD molecular subtypes were identified, exhibiting significant differences in prognosis and immune characteristics. A corresponding risk score model was established. Patients in the low-risk group showed better prognosis and enhanced immune activity, whereas the high-risk group displayed higher TMB and stemness scores. Drug sensitivity analysis revealed distinct responses to chemotherapeutic agents such as cisplatin and docetaxel between risk groups. Functional assays demonstrated that MTCH2 knockout significantly inhibited LUAD cell proliferation, migration, and invasion, and induced G0/G1 phase arrest, suggesting that MTCH2 may act as a potential adverse prognostic marker.

CONCLUSION: MD-related genes exhibit strong prognostic and immune subtyping value. The proposed risk model holds clinical potential, and MTCH2 may serve as a promising target for precision therapy in LUAD.

PMID:40510359 | PMC:PMC12159055 | DOI:10.3389/fimmu.2025.1585505

A deep learning framework for <em>in silico</em> screening of anticancer drugs at the single-cell level

Natl Sci Rev. 2024 Dec 10;12(2):nwae451. doi: 10.1093/nsr/nwae451. eCollection 2025 Feb.

ABSTRACT

Tumor heterogeneity plays a pivotal role in tumor progression and resistance to clinical treatment. Single-cell RNA sequencing (scRNA-seq) enables us to explore heterogeneity within a cell population and identify rare cell types, thereby improving our design of targeted therapeutic strategies. Here, we use a pan-cancer and pan-tissue single-cell transcriptional landscape to reveal heterogeneous expression patterns within malignant cells, precancerous cells, as well as cancer-associated stromal and endothelial cells. We introduce a deep learning framework named Shennong for in silico screening of anticancer drugs for targeting each of the landscape cell clusters. Utilizing Shennong, we could predict individual cell responses to pharmacologic compounds, evaluate drug candidates' tissue damaging effects, and investigate their corresponding action mechanisms. Prioritized compounds in Shennong's prediction results include FDA-approved drugs currently undergoing clinical trials for new indications, as well as drug candidates reporting anti-tumor activity. Furthermore, the tissue damaging effect prediction aligns with documented injuries and terminated discovery events. This robust and explainable framework has the potential to accelerate the drug discovery process and enhance the accuracy and efficiency of drug screening.

PMID:39872221 | PMC:PMC11771446 | DOI:10.1093/nsr/nwae451

Digital phenotyping from wearables using AI characterizes psychiatric disorders and identifies genetic associations

Complex disorders require precise strategies for their characterization. AI-based digital phenotypes from biosensors can be used to predict psychiatric disorders and identify GWAS loci.

Reprogramming tumour-associated macrophages to outcompete cancer cells

Nature, Published online: 28 June 2023; doi:10.1038/s41586-023-06256-5

In a mouse model of breast cancer, a low-protein diet induces engulfment activities and mTORC1 signalling in tumour-associated macrophages to suppress engulfment-dependent mTORC1 signalling in MYC-overexpressing cancer cells through cell competition, serving as an innate immune defence mechanism to slow tumour growth.
❌