❌

Reading view

Machine Learning-Driven Insights in Cancer Metabolomics: From Subtyping to Biomarker Discovery and Prognostic Modeling

Metabolites. 2025 Aug 1;15(8):514. doi: 10.3390/metabo15080514.

ABSTRACT

Cancer metabolic reprogramming plays a critical role in tumor progression and therapeutic resistance, underscoring the need for advanced analytical strategies. Metabolomics, leveraging mass spectrometry and nuclear magnetic resonance (NMR) spectroscopy, offers a comprehensive and functional readout of tumor biochemistry. By enabling both targeted metabolite quantification and untargeted profiling, metabolomics captures the dynamic metabolic alterations associated with cancer. The integration of metabolomics with machine learning (ML) approaches further enhances the interpretation of these complex, high-dimensional datasets, providing powerful insights into cancer biology from biomarker discovery to therapeutic targeting. This review systematically examines the transformative role of ML in cancer metabolomics. We discuss how various ML methodologies-including supervised algorithms (e.g., Support Vector Machine, Random Forest), unsupervised techniques (e.g., Principal Component Analysis, t-SNE), and deep learning frameworks-are advancing cancer research. Specifically, we highlight three major applications of ML-metabolomics integration: (1) cancer subtyping, exemplified by the use of Similarity Network Fusion (SNF) and LASSO regression to classify triple-negative breast cancer into subtypes with distinct survival outcomes; (2) biomarker discovery, where Random Forest and Partial Least Squares Discriminant Analysis (PLS-DA) models have achieved >90% accuracy in detecting breast and colorectal cancers through biofluid metabolomics; and (3) prognostic modeling, demonstrated by the identification of race-specific metabolic signatures in breast cancer and the prediction of clinical outcomes in lung and ovarian cancers. Beyond these areas, we explore applications across prostate, thyroid, and pancreatic cancers, where ML-driven metabolomics is contributing to earlier detection, improved risk stratification, and personalized treatment planning. We also address critical challenges, including issues of data quality (e.g., batch effects, missing values), model interpretability, and barriers to clinical translation. Emerging solutions, such as explainable artificial intelligence (XAI) approaches and standardized multi-omics integration pipelines, are discussed as pathways to overcome these hurdles. By synthesizing recent advances, this review illustrates how ML-enhanced metabolomics bridges the gap between fundamental cancer metabolism research and clinical application, offering new avenues for precision oncology through improved diagnosis, prognosis, and tailored therapeutic strategies.

PMID:40863133 | PMC:PMC12388062 | DOI:10.3390/metabo15080514

  •  

A multilevel genomic approach to uncover causal connections between CPFE and lung cancer subtypes: A two-sample Mendelian randomization study

Medicine (Baltimore). 2025 Aug 22;104(34):e44050. doi: 10.1097/MD.0000000000044050.

ABSTRACT

Combined pulmonary fibrosis and emphysema (CPFE) and lung cancer cast intertwined shadows, yet the molecular nexus binding them remains largely obscured. By integrating high-resolution transcriptomic landscapes, extensive genome-wide association resources, and a stratified Mendelian randomization (MR) framework, we distilled 809 differentially expressed genes and, in successive steps, confirmed their causal ties to squamous cell carcinoma, adenocarcinoma, and small cell lung cancer. The credibility of these associations was bolstered through 3 sequential validation tiers - eQTL-anchored MR, eQTL-anchored SMR, and pQTL-anchored MR analyses - each reinforcing the robustness of the signals. Within this constellation, CPPED1 emerged as a watchful sentinel that mitigates risk in squamous carcinoma, whereas CD300LF proved a formidable oncogenic catalyst in the small cell lineage. Collectively, these insights illuminate the heritable circuitry linking CPFE and lung cancer, chart avenues for proactive surveillance and precision therapeutics in vulnerable patients, and enrich the conceptual framework of the fibrosis-to-carcinoma transition, inviting deeper multi-omic synthesis and incisive mechanistic exploration.

PMID:40859573 | PMC:PMC12385043 | DOI:10.1097/MD.0000000000044050

  •  

RPN1 Is Associated With Immunosuppression in Pan-Cancer and Affects the Malignant Phenotype of Tumor

FASEB J. 2025 Aug 31;39(16):e70978. doi: 10.1096/fj.202500722RR.

ABSTRACT

Ribophorin 1 (RPN1), a key component of the oligosaccharyltransferase complex, is implicated in tumor progression through glycosylation-mediated pathways, yet its pan-cancer roles remain unexplored. This study presents a comprehensive multi-omics analysis of RPN1 across 33 cancers such as sarcoma (SARC), integrating genomic, transcriptomic, and proteomic data from TCGA, GTEx, and CPTAC. RPN1 was significantly overexpressed in 14 malignancies and correlated with advanced tumor stages and poor prognosis in glioblastoma (GBM), lower-grade glioma, SARC and hepatocellular carcinoma, validated in an independent glioma cohort (n = 151). Genomically, RPN1 amplification linked to homologous recombination deficiency and elevated tumor mutational burden, suggesting a role in genomic instability. Critically, multiplex immunofluorescence demonstrates RPN1 overexpression colocalizes with CD206+ M2 macrophages in tumor microenvironments, while in vitro coculture experiments confirm RPN1-dependent microglial recruitment and M2 polarization. RPN1 expression negatively correlates with CD8+ T cell infiltration and predicts resistance to chemotherapy (GBM, ovarian cancer) and immunotherapy (GBM, esophageal carcinoma), though it associates with PD-1 inhibitor sensitivity in bladder cancer. Functional validation shows RPN1 knockdown suppresses proliferation, migration, and invasion in GBM cells. Pathway enrichment connects RPN1 to endoplasmic reticulum stress, glycosylation, DNA repair, and immune checkpoint regulation. These findings position RPN1 as a multimodal oncogenic driver promoting genomic instability, immunosuppressive microenvironment remodeling, and context-dependent therapeutic vulnerabilities across cancers.

PMID:40857034 | DOI:10.1096/fj.202500722RR

  •  

The emerging role of microbiota in lung cancer: a new perspective on lung cancer development and treatment

Cell Oncol (Dordr). 2025 Aug 26. doi: 10.1007/s13402-025-01103-3. Online ahead of print.

ABSTRACT

Lung cancer remains the leading cause of cancer-related mortality worldwide, with limited treatment efficacy and frequent resistance to conventional therapies. Recent advances have uncovered the critical influence of the human microbiota-complex communities of bacteria, viruses, fungi, and other microorganisms-on lung cancer pathogenesis and therapeutic responses. This review synthesizes current knowledge on the compositional and functional roles of microbiota across multiple body sites, including the gut, lung, tumor microenvironment, circulation, and oral cavity, highlighting their contributions to tumor initiation, progression, metastasis, and immune regulation. We emphasize the bidirectional communication between microbial metabolites and host immune pathways, particularly the gut-lung axis, which modulates systemic and local antitumor immunity. Importantly, microbiota composition has been linked to differential responses and toxicities in chemotherapy, radiotherapy, targeted therapy, and immune checkpoint blockade. Microbiota-targeted interventions, such as probiotics, fecal microbiota transplantation, and selective antibiotics, show promising potential to enhance treatment efficacy and mitigate adverse effects. However, challenges remain in clinical translation due to interindividual microbiome variability, mechanistic complexities, and limited longitudinal data. Future research integrating multi-omics, microbial functional profiling, and controlled clinical trials is essential to harness the microbiome as a precision medicine tool in lung cancer management. This review provides a comprehensive overview of the emerging role of microbiota in lung cancer development and therapy, offering new perspectives for innovative therapeutic strategies.

PMID:40856929 | DOI:10.1007/s13402-025-01103-3

  •  

Systema: a framework for evaluating genetic perturbation response prediction beyond systematic variation

Nature Biotechnology, Published online: 25 August 2025; doi:10.1038/s41587-025-02777-8

An evaluation framework isolates perturbation-specific effects in perturbation datasets.
  •  

Refining treatment strategies for non-small cell lung cancer lacking actionable mutations: insights from multi-omics studies

Br J Cancer. 2025 Aug 23. doi: 10.1038/s41416-025-03139-6. Online ahead of print.

ABSTRACT

Non-small cell lung cancer (NSCLC) represents a heterogeneous group of malignancies characterised by diverse histological and molecular features. Some NSCLCs, particularly adenocarcinomas, harbour genomic alterations in receptor tyrosine kinases or downstream RAS/RAF signalling pathways, which are targets of effective therapies. NSCLCs lacking actionable genomic alterations often benefit from immune checkpoint inhibitors, though only a minority of patients achieve long-term survival. These tumours often carry alterations in tumour suppressor genes like TP53, KEAP1, STK11, or NF1, for which pharmacological strategies are still under investigation. This review explores emerging therapeutic opportunities unveiled by multi-omics studies in NSCLCs without actionable genomic alterations. Proteogenomic approaches-integrating genomic, transcriptomic and proteomic data-enable a comprehensive understanding of NSCLC molecular landscapes and signalling network dysregulation, helping to identify distinct tumour subtypes and potential therapeutic targets. These tumours exhibit alterations in cell cycle regulation, DNA repair, immune signalling, epigenetic modulation and metabolic and redox pathways. Although therapies targeting tumour suppressor genes like p53 remain highly anticipated, extending our understanding of the broader molecular landscape in these tumours may reveal novel vulnerabilities and inform the development of novel drugs or combination strategies. This could further advance precision oncology for NSCLC.

PMID:40849356 | DOI:10.1038/s41416-025-03139-6

  •  

Integrative genomic identification of therapeutic targets for pancreatic cancer

Cell Rep. 2025 Aug 21;44(9):116191. doi: 10.1016/j.celrep.2025.116191. Online ahead of print.

ABSTRACT

Pancreatic ductal adenocarcinoma (PDAC) is a deadly disease, and new therapeutic strategies are urgently needed. Here, we conduct an integrative, genome-scale examination of genetic dependencies and cell surface targets using CRISPR-Cas screening and multi-omic data, including single-nucleus and spatial transcriptomic data from patient tumors. We systematically identify clinically tractable and biomarker-linked PDAC dependencies, including CDS2 as a synthetic lethal target in cancer cells expressing signatures of epithelial-to-mesenchymal transition. We examine biomarkers and co-dependencies of the KRAS oncogene, defining gene expression signatures of sensitivity and resistance associated with response to pharmacological inhibition of KRAS. mRNA and protein profiling reveal cell surface protein-encoding genes with robust expression in patient tumors and minimal expression in non-malignant tissues. Furthermore, we define intratumoral and interpatient heterogeneity of target gene expression and identify orthogonal targets that suggest combinatorial strategies. Collectively, this work identifies multiple targets that may inform therapeutic strategies for patients with PDAC.

PMID:40848256 | DOI:10.1016/j.celrep.2025.116191

  •  

Targeting spermine metabolism to overcome immunotherapy resistance in pancreatic cancer

Nat Commun. 2025 Aug 22;16(1):7827. doi: 10.1038/s41467-025-63146-2.

ABSTRACT

While dysregulation of polyamine metabolism is frequently observed in cancer, it is unknown how polyamines alter the tumor microenvironment (TME) and contribute to therapeutic resistance. Analysis of polyamines in the plasma of pancreatic cancer patients reveals that spermine levels are significantly elevated and correlate with poor prognosis. Using a multi-omics approach, we identify Serpinb9 as a vulnerability in spermine metabolism in pancreatic cancer. Serpinb9, a serine protease inhibitor, directly interacts with spermine synthase (SMS), impeding its lysosome-mediated degradation and thereby augmenting spermine production and secretion. Mechanistically, the accumulation of spermine in the TME alters the metabolic landscape of immune cells, promoting CD8+ T cell dysfunction and pro-tumor polarization of macrophages, thus creating an immunosuppressive microenvironment. Small peptides that disrupt the Serpinb9-SMS interaction significantly enhance the efficacy of immune checkpoint blockade therapy. Together, our findings suggest that targeting spermine metabolism is a promising strategy to improve pancreatic cancer immunotherapy.

PMID:40846845 | PMC:PMC12373741 | DOI:10.1038/s41467-025-63146-2

  •  

Simplifying protein engineering with deep learning

When it comes to deep learning for protein engineering, there is strength in simplicity. In this issue of Cell, with thoughtful deployment of existing fixed-backbone sequence design models, Caixia Gao and colleagues engineer diverse genome editing systems with improved functionality, enabling powerful capabilities in fine-grained and large-scale genome editing as demonstrated through strong experimental validation.
  •  

Interferons in health and disease

The cytokine messenger proteins known as interferons are central to protective immune responses against infections, but they are also involved in inflammatory and autoimmune diseases. This review maps out the balance of factors that governs the many roles of interferons in animal biology.
  •  

Human interpretable grammar encodes multicellular systems biology models to democratize virtual cell laboratories

We developed a plain text modeling language—a cell behavior hypothesis grammar—to easily build virtual cell models and connect them to data, helping scientists to unlock the hidden dynamics of tissues. We provide examples showing how to use them in virtual experiments exploring how cancer responds to the cells in its environment and how the brain forms layers in development.
  •  

Circulating tumour cells & circulating tumour DNA in patients with resectable colorectal liver metastases (MIRACLE): a prospective, observational biomarker study

EClinicalMedicine. 2025 Aug 12;87:103406. doi: 10.1016/j.eclinm.2025.103406. eCollection 2025 Sep.

ABSTRACT

BACKGROUND: Recurrence risk after curative surgery for colorectal liver metastases (CRLM) remains high, underlining the need to identify prognostic markers enabling more individualised treatment approaches.

METHODS: In the MIRACLE, a prospective, observational biomarker study, a total of 188 patients with isolated, resectable CRLM without (neo)adjuvant chemotherapy were included between October 2015 and December 2021. Blood samples were collected before surgery (baseline) and three weeks after surgery. The primary objective was to assess the potential association between postoperative circulating tumour DNA (ctDNA) detection and recurrence of disease for patients with resectable CRLM within one year after resection. The secondary objective was the association between recurrence of disease within one year and detection of circulating tumour cells (CTCs). Baseline ctDNA was measured by next generation sequencing using a targeted panel (Oncomine Colon cell-free DNA assay) and postoperatively by digital PCR on genetic variants found preoperatively with the Oncomine panel. CTCs were enumerated using the FDA-approved CellSearch system.

FINDINGS: ctDNA was detected in 117/187 patients (63%) at baseline, and 28/104 evaluable patients (27%) still had detectable ctDNA postoperatively. CTC enumeration resulted in positivity for 37/183 patients (20%) at baseline and 14/158 patients (9%) postoperatively. No association was found between 1-year recurrence-free survival (RFS) and the presence of CTCs or ctDNA at baseline. In contrast, patients with postoperative undetectable ctDNA had a significantly improved 1-year RFS compared to patients with postoperative ctDNA (54% [95% CI 44%-67%] vs. 25% [95% CI 13%-47%], log-rank p = 0.0011). Similarly, patients with postoperative detectable CTCs had a significantly shorter 1-year RFS compared to patients without postoperative CTCs (15% [95% CI 4%-55%] vs. 53% [95% CI 45%-62%], log-rank p 0.0004). Also in multivariable analysis, detectable ctDNA and CTCs after surgery remained independently associated with a shorter 1-year RFS (HR 2.35; 95% CI 1.34-4.11; p = 0.0028 and HR 2.98; 95% CI 1.56-5.71; p = 0.0010, respectively).

INTERPRETATION: This is the first study conducted in patients with resectable CRLM without (neo)adjuvant chemotherapy, which demonstrates the impact of postoperative detectable circulating tumour load on 1-year RFS. Postoperative ctDNA and CTC detection both represent strong, independent predictors for a shorter RFS after local treatment, as opposed to preoperative detection.

FUNDING: This work was supported by KWF Kankerbestrijding (Dutch Cancer Society, EMCR 2014-6340).

PMID:40838198 | PMC:PMC12361997 | DOI:10.1016/j.eclinm.2025.103406

  •  

Genetic and epigenetic dysregulation of CR1 is associated with catastrophic antiphospholipid syndrome

Ann Rheum Dis. 2025 Aug 20:S0003-4967(25)04249-9. doi: 10.1016/j.ard.2025.07.016. Online ahead of print.

ABSTRACT

OBJECTIVES: Catastrophic antiphospholipid syndrome (CAPS) is a complement-driven thrombotic disorder, characterised by widespread thrombosis and multiorgan failure. We identified rare germline variants including complement receptor 1 (CR1) in 50% of patients with CAPS. Here, we define CR1 dysregulation mechanisms (genetic/epigenetic) underlying complement-mediated thrombosis in CAPS and support C5 inhibition as a potential therapy.

METHODS: We quantified CR1 expression by flow cytometry across haematopoietic cell types. CRISPR/Cas9 genome editing of TF-1 (erythroleukaemia) cells was performed to generate CR1 'knock-out' and 'knock-in' lines with patient-specific CR1 variants. Multiomics analysis was performed to investigate the role of methylation in patients with reduced CR1 expression. Functional impact of low CR1 was assessed by complement-mediated cell killing using modified Ham assay, cell-bound complement degradation products through flow cytometry, and circulatory immune complexes in serum samples through ELISA.

RESULTS: CR1 expression in erythrocytes was markedly reduced on CAPS erythrocytes (n = 9, 21.80%) compared to healthy controls (HCs; n = 35, 84.04%), with promoter hypermethylation emerging as a plausible epigenetic mechanism for CR1 downregulation. Novel germline variant (CR1-V2125L; rs202148801) mitigated CR1 expression and increased complement-mediated cell death of knock-in cell lines. Erythrocytes from the patient with the CR1-V2125L variant had low CR1 expression. Levels of circulating immune complexes, which are bound and cleared by CR1 on erythrocytes, were higher in acute CAPS (n = 3, 25.55 µg Eq/mL) than HCs (n = 3, 7.445 µg Eq/mL). Five patients were treated with C5 inhibition which mitigated thrombosis.

CONCLUSIONS: Genetic or epigenetic-mediated CR1 deficiency is a potential hallmark of CAPS and predicts response to C5 inhibition.

PMID:40841298 | DOI:10.1016/j.ard.2025.07.016

  •  

Liquid biopsy in lung cancer

Breathe (Sheff). 2025 Aug 19;21(3):250051. doi: 10.1183/20734735.0051-2025. eCollection 2025 Jul.

ABSTRACT

Lung cancer is the leading cause of cancer-related mortality worldwide, with nonsmall cell lung cancer (NSCLC) accounting for the majority of cases. Despite advancements in therapeutics, outcomes remain poor due to late-stage diagnoses and the molecular complexity of the disease. Liquid biopsy, a minimally invasive diagnostic approach, has emerged as a potentially transformative tool in lung cancer. The detection of tumour-derived biomarkers, such as circulating-tumour DNA, circulating tumour cells and exosomes, can be analysed for molecular profiling, early detection and monitoring of disease progression. There have been significant advancements of liquid biopsy technologies, such as next-generation sequencing and droplet digital PCR, that identify actionable mutations, detect resistance mechanisms and improve therapeutic outcomes. While there are still challenges like detecting early-stage disease and the risk of false positives, the combination of multi-omics data and artificial intelligence has the potential for more personalised and precise cancer treatments. Liquid biopsy represents a paradigm shift in the early detection and personalised treatment of lung cancer, offering significant potential to improve patient outcomes.

PMID:40837417 | PMC:PMC12362143 | DOI:10.1183/20734735.0051-2025

  •  

In a first, Google has released data on how much energy an AI prompt uses

Google has just released a technical report detailing how much energy its Gemini apps use for each query. In total, the median prompt—one that falls in the middle of the range of energy demand—consumes 0.24 watt-hours of electricity, the equivalent of running a standard microwave for about one second. The company also provided average estimates for the water consumption and carbon emissions associated with a text prompt to Gemini.

It’s the most transparent estimate yet from a Big Tech company with a popular AI product, and the report includes detailed information about how the company calculated its final estimate. As AI has become more widely adopted, there’s been a growing effort to understand its energy use. But public efforts to directly measure the energy used by AI have been hampered by a lack of full access to the operations of a major tech company. 

Earlier this year, MIT Technology Review published a comprehensive series on AI and energy, at which time none of the major AI companies would reveal their per-prompt energy usage. Google’s new publication, at last, allows for a peek behind the curtain that researchers and analysts have long hoped for.

The study focuses on a broad look at energy demand, including the power used not only by the AI chips that run models but also by all the other infrastructure needed to support that hardware. 

“We wanted to be quite comprehensive in all the things we included,” said Jeff Dean, Google’s chief scientist, in an exclusive interview with MIT Technology Review about the new report.

That’s significant, because in this measurement, the AI chips—in this case, Google’s custom TPUs, the company’s proprietary equivalent of GPUs—account for just 58% of the total electricity demand of 0.24 watt-hours. 

Another large portion of the energy is used by equipment needed to support AI-specific hardware: The host machine’s CPU and memory account for another 25% of the total energy used. There’s also backup equipment needed in case something fails—these idle machines account for 10% of the total. The final 8% is from overhead associated with running a data center, including cooling and power conversion. 

This sort of report shows the value of industry input to energy and AI research, says Mosharaf Chowdhury, a professor at the University of Michigan and one of the heads of the ML.Energy leaderboard, which tracks energy consumption of AI models. 

Estimates like Google’s are generally something that only companies can produce, because they run at a larger scale than researchers are able to and have access to behind-the-scenes information. “I think this will be a keystone piece in the AI energy field,” says Jae-Won Chung, a PhD candidate at the University of Michigan and another leader of the ML.Energy effort. “It’s the most comprehensive analysis so far.”

Google’s figure, however, is not representative of all queries submitted to Gemini: The company handles a huge variety of requests, and this estimate is calculated from a median energy demand, one that falls in the middle of the range of possible queries.

So some Gemini prompts use much more energy than this: Dean gives the example of feeding dozens of books into Gemini and asking it to produce a detailed synopsis of their content. “That’s the kind of thing that will probably take more energy than the median prompt,” he says. Using a reasoning model could also have a higher associated energy demand because these models take more steps before producing an answer.

This report was also strictly limited to text prompts, so it doesn’t represent what’s needed to generate an image or a video. (Other analyses, including one in MIT Technology Review’s Power Hungry series earlier this year, show that these tasks can require much more energy.)

The report also finds that the total energy used to field a Gemini query has fallen dramatically over time. The median Gemini prompt used 33 times more energy in May 2024 than it did in May 2025, according to Google. The company points to advancements in its models and other software optimizations for the improvements.  

Google also estimates the greenhouse-gas emissions associated with the median prompt, which they put at 0.03 grams of carbon dioxide. To get to this number, the company multiplied the total energy used to respond to a prompt by the average emissions per unit of electricity.

Rather than using an emissions estimate based on the US grid average, or the average of the grids where Google operates, the company instead uses a market-based estimate, which takes into account electricity purchases that the company makes from clean energy projects. The company has signed agreements to buy over 22 gigawatts of power from sources including solar, wind, geothermal, and advanced nuclear projects since 2010. Because of those purchases, Google’s emissions per unit of electricity on paper are roughly one-third of those on the average grid where it operates.

AI data centers also consume water for cooling, and Google estimates that each prompt consumes 0.26 milliliters of water, or about five drops. 

The goal of this work was to provide users a window into the energy use of their interactions with AI, Dean says. 

“People are using [AI tools] for all kinds of things, and they shouldn’t have major concerns about the energy usage or the water usage of Gemini models, because in our actual measurements, what we were able to show was that it’s actually equivalent to things you do without even thinking about it on a daily basis,” he says, “like watching a few seconds of TV or consuming five drops of water.”

The publication greatly expands what’s known about AI’s resource usage. It follows recent increasing pressure on companies to release more information about the energy toll of the technology. “I’m really happy that they put this out,” says Sasha Luccioni, an AI and climate researcher at Hugging Face. “People want to know what the cost is.”

This estimate and the supporting report contain more public information than has been available before, and it’s helpful to get more information about AI use in real life, at scale, by a major company, Luccioni adds. However, there are still details that the company isn’t sharing in this report. One major question mark is the total number of queries that Gemini gets each day, which would allow estimates of the AI tool’s total energy demand. 

And ultimately, it’s still the company deciding what details to share, and when and how. “We’ve been trying to push for a standardized AI energy score,” Luccioni says, a standard for AI similar to the Energy Star rating for appliances. “This is not a replacement or proxy for standardized comparisons.”

  •  

Discovery of a novel potent tubulin inhibitor through virtual screening and target validation for cancer chemotherapy

Cell Death Discovery, Published online: 19 August 2025; doi:10.1038/s41420-025-02679-3

Discovery of a novel potent tubulin inhibitor through virtual screening and target validation for cancer chemotherapy
  •  
❌