❌

Normal view

Toward the best generalizable performance of machine learning in modeling omic and clinical data

17 October 2025 at 18:00

Lab Invest. 2025 Oct 15:104253. doi: 10.1016/j.labinv.2025.104253. Online ahead of print.

ABSTRACT

There are often performance differences between intra-dataset and cross-dataset tests in machine learning (ML) modeling. However, reducing these differences may reduce ML performances. It is thus a challenging dilemma for developing models that excel in intra-dataset testing and are generalizable to cross-dataset testing. Therefore, we aimed to understand and improve performance and generalizability of ML in intra-dataset and cross-dataset testing. We evaluated 4,200 ML models of classifying lung adenocarcinoma (LUAD) deaths using the The Cancer Genome Atlas (TCGA, n=286) and Oncogenomic-Singapore (OncoSG, n=167) datasets, and 1,680 models of classifying glioblastoma deaths using TCGA (n=151) and Clinical Proteomic Tumor Analysis Consortium (CPTAC, n=97) datasets. After examining performance distributions of these ML models, we applied a dual analytical framework, including statistical analyses and SHapley Additive exPlanations-based meta-analysis, to quantify factors' importance and trace model success back to design principles. We also developed a framework to identify the best generalizable model. Strikingly, Jarque-Bera test revealed significant deviations of model performances from normality in both cancer types and testing contexts. Simple linear models with sparse feature sets consistently dominated in LUAD experiments, whereas non-linear models dominated in glioblastoma ones, suggesting that the best modeling strategy appears cancer-type/disease dependent. Importantly, both robust Analysis of Variance (ANOVA) and Kruskal-Wallis tests consistently identified differentially expressed genes as one of the most influential factors in both cancer types. The proposed multi-criteria framework successfully identified the model that achieved both the best cross-dataset performance and similar intra-dataset performance. In summary, ML performance distributions significantly deviated from normality, which motivates using both robust parametric and non-parametric statistical tests. We quantified and provided possible exploitability on the factors associated with cross-dataset performances and generalizability of ML models in two cancer types. A multi-criteria framework was developed and validated to identify the models that are accurate and consistently robust cross datasets.

PMID:41106592 | DOI:10.1016/j.labinv.2025.104253

Multi-omics analyses inform mechanisms of immunotherapy response in pancreatic cancer

Front Immunol. 2025 Oct 2;16:1673098. doi: 10.3389/fimmu.2025.1673098. eCollection 2025.

ABSTRACT

INTRODUCTION: Pancreatic ductal adenocarcinoma (PDAC) continues to exhibit resistance to immunotherapy. In this study, we evaluated the efficacy of combining immunotherapy with chemotherapy for the treatment of advanced pancreatic cancer. Additionally, we employed a multimodal analytical approach to elucidate the immune landscape and conduct transcriptomic profiling in PDAC.

METHODS: A retrospective analysis was conducted on the clinical data of 52 patients diagnosed with advanced PDAC who underwent a combined treatment regimen of immunotherapy and chemotherapy. The study evaluated the objective response rate (ORR), disease control rate (DCR), and progression-free survival (PFS). To characterize the immune landscape in treatment-naive pancreatic ductal adenocarcinoma (PDAC) tumors and in the systemic circulation, flow cytometry, multiplex immunohistochemistry (mIHC), and whole transcriptome sequencing were employed.

RESULTS: The study reported an ORR of 32.7%, a DCR of 67.3%, and a 6-month PFS rate of 38.5%, with a median PFS of 5.5 months. Patients treated with a combination of immunotherapy and gemcitabine achieved the longest PFS. The first-line treatment cohort exhibited a significantly higher DCR (79.3% vs. 52.2%, P = 0.038) and a longer median PFS (6.6 vs. 3.5 months, P = 0.032) compared to the second-line treatment cohort. The efficacy of treatment varied depending on the drug combinations used. Flow cytometry analysis revealed a greater frequency of CD45- CD64+ cells in the peripheral blood of patients with progressive disease (PD) compared to those with a partial response (PR). Multiplex immunofluorescence (MIF) analysis indicated an increased intratumoral infiltration of CD8+ T cells and CD137+ CD8+ T cells in patients with PR. Whole transcriptome sequencing (WTSS) identified key genes involved in immune regulation, signal transduction, and digestive function. Hemopexin (HPX) and regulatory factor X-associated protein (RFXAP) were upregulated in PR patients and showed a positive correlation with survival, whereas Interleukin-6 (IL-6) expression was linked to poor prognosis.

CONCLUSIONS: These findings indicate that immunochemotherapy shows potential for the treatment of advanced PDAC. Our study elucidates the immune landscape associated with PDAC and provides critical insights for the identification of prospective therapeutic targets, which could guide the development of innovative combination immunotherapy strategies.

PMID:41112307 | PMC:PMC12528169 | DOI:10.3389/fimmu.2025.1673098

Integrative Transcriptomic and Metabolomic Analysis Reveals Aberrant Glycosylation as a Hallmark of Lung Adenocarcinoma

17 October 2025 at 18:00

OMICS. 2025 Oct 16. doi: 10.1177/15578100251387518. Online ahead of print.

ABSTRACT

Lung adenocarcinoma (LUAD) remains the most common subtype of lung cancer, characterized by high heterogeneity and poor survival outcomes. Although transcriptomic and metabolomic alterations have been individually studied, integrated multi-omics analyses are needed to uncover the convergent pathways that drive tumor progression. Differentially expressed genes (DEGs) were identified from the GSE229253 transcriptomic dataset comprising LUAD tumor and adjacent normal tissues, while significantly altered metabolites were obtained from the Lung Cancer Metabolome Database. The top 10 DEGs and metabolites were analyzed using the search tool for interacting chemicals (STITCH) to construct gene-metabolite networks, and Integrated Molecular Pathway Level Analysis (IMPaLA) was employed for integrated pathway enrichment to identify overlapping molecular processes. Transcriptomic profiling revealed 973 DEGs (410 upregulated and 563 downregulated), and metabolomic analysis identified significant alterations in metabolites linked to redox balance, amino acid derivatives, and nucleotide metabolism. Integration through STITCH generated a network of 16 nodes and 9 edges, highlighting gene-metabolite associations of probable biological relevance. Joint pathway enrichment analysis using IMPaLA consistently identified glycosylation-related pathways, particularly O-linked glycosylation of mucins, as major axes of convergence between transcriptomic and metabolomic alterations in LUAD (joint p = 0.00129-0.00434). Several genes (B3GNT6, FEZF1-AS1, and LCAL1) and metabolites (isoleucylleucine, leucylleucine, and isoleucylvaline) are probable novel candidates, warranting further investigation. These findings provide systems-level evidence that aberrant glycosylation is likely a central hallmark of LUAD, underscore the potential of glycosylation pathways as biomarkers and therapeutic targets, and demonstrate the utility of cross-omics approaches to unpack the molecular complexity of lung cancer.

PMID:41103242 | DOI:10.1177/15578100251387518

Are cancer surgeries removing the body’s secret weapon against cancer?

20 October 2025 at 19:48
Scientists have found that preserving lymph nodes during cancer surgery could dramatically improve how patients respond to immunotherapy. The research shows that lymph nodes are essential for training and sustaining cancer-fighting T cells. Removing them may unintentionally weaken the immune response, while keeping them intact could help unlock stronger, longer-lasting treatments.
  • ✇Nature Medicine
  • <b>Creativity keeps the brain young</b> Karen O’Leary
    Nature Medicine, Published online: 20 October 2025; doi:10.1038/d41591-025-00063-3A study shows that creativity enhances brain health by improving connectivity in age-vulnerable regions, highlighting the importance of supporting creative activities in public health strategies.
     

Capacity to Invest Effort as a Predictor of Preference for Digital Mental Health Interventions Over Psychotherapy: Cross-Sectional Study Using an Ecological Digital Screening Tool

Background: Research typically shows a higher preference for professionally-led face-to-face mental health interventions over digital ones. It remains unclear in which circumstances digital self-help tools are preferred. To address this gap, it is important to examine user characteristics that may help predict when digital interventions are more desirable, ultimately guiding their design to enhance engagement and appeal. Objective: To examine how distress severity and capacity to invest effort relate to intervention preferences, using an ecological assessment of individuals who seek to receive feedback on their mental health. Methods: A comprehensive digital mental health screening tool providing automated feedback was developed and advertised on social media. The sample comprised 684 adult participants aged 18-82 who opted to complete the screening to receive feedback on their mental health state. Participants completed questionnaires measuring general psychological distress, depression, generalized anxiety and demographics. Kessler Psychological Distress Scale–6 was used as the primary measure for distress. Participants were also presented with questions measuring capacity to invest effort and preferences for a professional vs digital self-help tools and for psychotherapy vs a mobile application. The effectiveness of distress, capacity to invest effort, and background characteristics in predicting preferences (a professional vs digital self-help tools; psychotherapy vs a mobile application) was examined using hierarchical linear regressions. The distributions of dichotomized preferences were plotted against distress and capacity to invest effort for transparent visualization. Results: A hierarchical linear regression found that distress, capacity to invest, and currently being in psychotherapy significantly predicted preference for a professional vs digital self-help tools. Distress (β=.25, 95% CI .18 to .32, P<.001 and capacity to invest effort ci .16 .30 p were the strongest predictors with similar effect size. model explained of variance in preference uniquely contributing most distressed participants low preferred digital self-help tools whereas high favored a professional. results obtained when using phq-4 as an alternative distress measure. remained significant .10 .26 predicting for psychotherapy vs mobile application while was not .05 conclusions: this study highlights that interventions is driven by reduced intervention. attempts reduce mental health treatment gap through should focus on optimizing elicited users improve desirability engagement.>
  • ✇STAT
  • STAT+: Duke data scientist launches startup to help hospitals adopt AI Casey Ross
    Mark Sendak was getting tired of seeing the toil of so many colleagues go to waste. At Duke University, he was part of a team of data scientists and engineers who built artificial intelligence tools to help make better health care decisions, and to more effectively treat patients with serious and life-threatening conditions.  But even when one of their inventions appeared to help patients and generated positive results in scientific studies, it never gained uptake beyond Duke’s walls. Pati
     

STAT+: Duke data scientist launches startup to help hospitals adopt AI

20 October 2025 at 16:30

Mark Sendak was getting tired of seeing the toil of so many colleagues go to waste.

At Duke University, he was part of a team of data scientists and engineers who built artificial intelligence tools to help make better health care decisions, and to more effectively treat patients with serious and life-threatening conditions. 

But even when one of their inventions appeared to help patients and generated positive results in scientific studies, it never gained uptake beyond Duke’s walls. Patients and doctors in other health systems didn’t get the opportunity to benefit.

Continue to STAT+ to read the full story…

© Courtesy Vega Health

Comprehensive bioinformatics analysis of omics data to reveal molecular mechanisms and biomarkers in multiple cancers

In Silico Pharmacol. 2025 Oct 17;13(3):154. doi: 10.1007/s40203-025-00440-3. eCollection 2025.

ABSTRACT

Breast, ovarian, lung, cervical, and colorectal cancers are among the most prevalent malignancies affecting women worldwide. This study aimed to elucidate the common molecular mechanisms of tumorigenesis and identify potential biomarkers using an integrative bioinformatics and network-based approach. Integrative profiling of five microarray datasets identified 66 differentially expressed genes (DEGs) that are common across five cancer types. Gene ontology and KEGG pathway analyses of common DEGs were performed using the DAVID database. The cell cycle processes were the most enriched functions, and oocyte meiosis, oocyte maturation, the p53 signaling pathway, cancer pathways, and cellular senescence were the most important pathways identified. Protein-protein interaction (PPI) networks for the DEGs were constructed using the STRING database, and the resulting networks were visualized in Cytoscape. Through PPI network analysis, ten hub genes were identified, and subsequent survival analysis confirmed that CHEK1, DLGAP5, CCNB2, and CCNA2 are significantly associated with poor patient survivability, establishing them as common biomarkers across multiple cancer types. Subsequently, ten transcription factors (TFs) and ten post-transcriptional regulators were identified through the assessment of regulatory networks involving TFs-DEGs and miRNAs-DEGs. Finally, drug-gene association analysis from the GSCA library was used to anticipate drug-like compounds using the drug repurposing approach. Overall, this comprehensive investigation holds promise for future in vitro and in vivo studies, offering a molecular foundation for the diagnosis, prognosis, and treatment of malignant cancers.

SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at 10.1007/s40203-025-00440-3.

PMID:41113171 | PMC:PMC12534660 | DOI:10.1007/s40203-025-00440-3

Pan-Cancer Analyses of Shared and Distinct Gene Expression in 17 Cancers: Rethinking Cancer Classification and Moving Beyond "One Drug, One Disease" Paradigm of Pharmaceutical Innovation

OMICS. 2025 Oct 17. doi: 10.1177/15578100251387873. Online ahead of print.

ABSTRACT

Cancer is a disease with heterogenous molecular signatures that ought to be unpacked to achieve the overarching aim of precision oncology. A pan-cancer omics approach provides a systems science framework to explore shared and distinct mechanisms across cancers. We report here pan-cancer analyses of gene expression data from 17 cancers, for example, adrenocortical cancer, lung cancer, kidney cancer, and colorectal cancer, and 26 tissue types, using public datasets to construct disease-specific transcriptional networks. Using the hypergeometric test, 1005 microRNAs (miRNAs), 314 transcription factors (TFs), and 332 receptors were identified as regulatory molecules interacting with differentially expressed genes. Kyoto Encyclopedia of Genes and Genomes pathway analysis was performed to explore their functional roles. Accordingly, we found miR-124-3p, miR-6799-5p, and miR-7106-5p as common miRNAs; Specificity Protein 1 (SP1), RELA Proto-Oncogene, NF-κB Subunit (RELA), and Nuclear Factor Kappa B Subunit 1 (NFKB1) as shared TFs; Cyclin-Dependent Kinase 2 (CDK2), Histone Deacetylase 1 (HDAC1), and ABL Proto-Oncogene 1, Non-Receptor Tyrosine Kinase (ABL1) as common receptors; and pathways in cancer, PI3K-Akt signaling, and p53 signaling as commonly enriched. Survival analysis in an independent dataset confirmed these findings: SP1 and NFKB1 were significant in 9 cancers, RELA in 6, whereas CDK2, HDAC1, and ABL1 were significant in 11, 10, and 10 cancers, respectively, out of the 17 cancers researched herein. In conclusion, these findings provide system-level insights on tumor heterogeneity and inform future cancer classification, for example, according to shared and distinct molecular signatures and development of therapies that might prove effective across several cancers. We underline that unpacking molecular signatures across multiple cancers also offers new prospects to move beyond the "One Drug, One Disease" paradigm of pharmaceutical innovation.

PMID:41111411 | DOI:10.1177/15578100251387873

Alternatives to animal testing are the future — it’s time that journals, funders and scientists embrace them

Nature, Published online: 20 October 2025; doi:10.1038/d41586-025-03344-6

Biomedical research techniques that don’t involve the use of animals are gaining momentum, but those using innovative approaches still face resistance from some quarters.

Circulating tumor DNA in Non-Viral head and neck squamous cell Carcinoma: A systematic review and Meta-Analysis

Oral Oncol. 2025 Nov;170:107760. doi: 10.1016/j.oraloncology.2025.107760. Epub 2025 Oct 17.

ABSTRACT

Non-viral head and neck squamous cell carcinoma (HNSCC) has poor survival and high recurrence rates. Circulating tumor DNA (ctDNA) is a promising biomarker for understanding tumor biology, assessing treatment response, and monitoring disease progression. While extensively studied in virally mediated HNSCC, its role in non-viral HNSCC remains underexplored. This systematic review and meta-analysis consolidates evidence on the diagnostic, prognostic, and therapeutic value of ctDNA in non-viral HNSCC. A systematic search across Medline, PubMed, Embase, and the Cochrane Library identified 1,915 records, of which 47 were included. Data extraction followed PRISMA guidelines, with overall survival (OS), progression-free survival (PFS), and recurrence-free survival (RFS), pooled as hazard ratios (HRs) with 95% confidence intervals (CIs) using a fixed-effect model. Among 3,574 patients, the most common tumor sites were the oral cavity (35 %) and oropharynx (22 %), with the majority presenting with stage IVA/IVB disease (29 %). Pre-treatment ctDNA detection rates ranged from 50 % to 100 % (median: 83 %), while post-treatment detection rates varied between 28 % and 100 % (median: 48 %). ctDNA detected recurrence in 80 % of patients, with a median lead time of 4.6 months. ctDNA detection was significantly associated with worse OS (HR 10.26, 95 % CI 3.58-29.40; P < 0.0001). Residual ctDNA was strongly correlated with worse PFS (HR 7.32, 95 % CI 4.17-12.86; P < 0.00001) and RFS (HR 7.33, 95 % CI 2.75-19.58; P < 0.0001). ctDNA holds potential for improving diagnostic accuracy, monitoring progression, and predicting survival outcomes in non-viral HNSCC. However, further large-scale studies and standardized guidelines are needed for validation and clinical implementation.

PMID:41108912 | DOI:10.1016/j.oraloncology.2025.107760

  • ✇InfoQ
  • Article: A Plan-Do-Check-Act Framework for AI Code Generation Ken Judy
    AI code generation tools promise faster development but often create quality issues, integration problems, and delivery delays. A structured Plan-Do-Check-Act cycle can maintain code quality while leveraging AI capabilities. Through working agreements, structured prompts, and continuous retrospection, it asserts accountability over code while guiding AI to produce tested, maintainable software. By Ken Judy
     

Article: A Plan-Do-Check-Act Framework for AI Code Generation

20 October 2025 at 19:00

AI code generation tools promise faster development but often create quality issues, integration problems, and delivery delays. A structured Plan-Do-Check-Act cycle can maintain code quality while leveraging AI capabilities. Through working agreements, structured prompts, and continuous retrospection, it asserts accountability over code while guiding AI to produce tested, maintainable software.

By Ken Judy
❌