❌

Normal view

Young Adults’ Interactions With Food and Nutrition Content on Social Media and Implications for Intervention Design: Semistructured Interview Study

Background: Young adults increasingly rely on social media for nutrition information. However, little is known about (1) which types of eating-related content they actively engage with and why, and (2) how they interpret, evaluate, and incorporate this content into their everyday food choices and health behaviors. Objective: This qualitative study explored how UK young adults (aged 18-25 years) interact with food and nutrition content across social media platforms to inform the design of future social media interventions. Methods: Semistructured online interviews, guided by the Capability, Opportunity, Motivation–Behavior (COM-B) model, were conducted with 25 active social media users (18/25, 72% women, mean age 22.2, SD 1.9 years, ethnically diverse) in the United Kingdom between August and October 2024. The study design was informed by patient and public involvement to ensure relevance and acceptability. Data were analyzed using reflexive thematic analysis. To guide intervention development, key findings (coded as barriers and facilitators) were systematically mapped to the Theoretical Domains Framework, and the COM-B. Ethics approval was obtained from the University of Cambridge (24.368). Results: Five key themes were identified: (1) evolving engagement patterns (passive scrolling to active interaction and mixed feelings on algorithmic control), (2) conflicted information seeking (frustration with contradictory advice, varied strategies to assess credibility), (3) multifaceted behavioral impact (simultaneous positive impacts such as cooking inspiration and negative impacts such as restrictive eating triggers), (4) shifting goals (a movement from appearance-focused to health-centered goals; yet, vulnerability to body-image issues), and (5) intervention preferences (demand for credible professionals, customizable content, and privacy protection). Participants demonstrated a reactive learning process, developing “digital nutrition literacy” often after negative experiences. Social influences were identified as the most frequently cited domain (mapped to TDF [theoretical domains framework]/COM-B) shaping interactions with social media content. Conclusions: This study challenges assumptions of passive social media consumption, showing that young adults actively develop protective strategies yet remain vulnerable to misinformation. Digital interventions should leverage user agency and address diverse perceptions through customizable, credible content delivered with privacy and emotionally safe messaging. The COM-B and TDF mapping provide specific, evidence-based behavioral targets, particularly within the domain of Social Opportunity and Reflective Motivation, to guide the development of effective eHealth interventions.
  • ✇STAT
  • STAT+: States looking to regulate use of chatbots Mario Aguilar
    You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday. Good morning health tech readers! Today, a deep dive into why America’s most powerful health insurer is looking more and more like a technology company. Continue to STAT+ to read the full story…
     

STAT+: States looking to regulate use of chatbots

7 April 2026 at 23:45

You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday.

Good morning health tech readers!

Today, a deep dive into why America’s most powerful health insurer is looking more and more like a technology company. 

Continue to STAT+ to read the full story…

© Tim Gruber for STAT

A Multimodal Foundation Model of Spatial Transcriptomics and Histology for Biological Discovery and Clinical Prediction

arXiv:2604.03630v1 Announce Type: new Abstract: Spatial transcriptomics (ST) enables gene expression mapping within anatomical context but remains costly and low-throughput. Hematoxylin and eosin (H\&E) staining offers rich morphology yet lacks molecular resolution. We present \textbf{\ours} (\textbf{S}patial \textbf{T}ranscriptomics and hist\textbf{O}logy \textbf{R}epresentation \textbf{M}odel), a foundation model trained on 1.2 million spatially resolved transcriptomic profiles with matched histology across 18 organs. Using a hierarchical architecture integrating morphological features, gene expression, and spatial context, STORM bridges imaging and omics through robust molecular--morphological representations. STORM enhances spatial domain discovery, producing biologically coherent tissue maps, and outperforms existing methods in predicting spatial gene expression from H\&E images across 11 tumor types. The model is platform-agnostic, performing consistently across Visium, Xenium, Visium HD, and CosMx. Applied to 23 independent cohorts comprising 7,245 patients, STORM significantly improves immunotherapy response prediction and prognostication over established biomarkers, providing a scalable framework for spatially informed discovery and clinical precision medicine.

PanLUNA: An Efficient and Robust Query-Unified Multimodal Model for Edge Biosignal Intelligence

arXiv:2604.04297v1 Announce Type: new Abstract: Physiological foundation models (FMs) have shown promise for biosignal representation learning, yet most remain confined to a single modality such as EEG, ECG, or PPG, largely because paired multimodal datasets are scarce. In this paper, we present PanLUNA, a compact 5.4M-parameter pan-modal FM that jointly processes EEG, ECG, and PPG within a single shared encoder. Extending LUNA's channel-unification module, PanLUNA treats multimodal channels as entries in a unified query set augmented with sensor-type embeddings, enabling efficient cross-modal early fusion while remaining inherently robust to missing modalities at inference time. Despite its small footprint, PanLUNA matches or exceeds models up to 57$\times$ larger: 81.21% balanced accuracy on TUAB abnormal EEG detection and state-of-the-art 0.7416 balanced accuracy on HMC multimodal sleep staging. Quantization-aware training with INT8 weights recovers $\geq$96% of full-precision performance, and deployment on the GAP9 ultra-low-power RISC-V microcontroller for wearables achieves 325.6 ms latency and 18.8 mJ per 10-second, 12-lead ECG inference, and 1.206 s latency at 68.65 mJ for multimodal 5-channel sleep staging over 30-second epochs.

Is your AI Model Accurate Enough? The Difficult Choices Behind Rigorous AI Development and the EU AI Act

arXiv:2604.03254v1 Announce Type: cross Abstract: Technical and legal debates frequently suggest that "accuracy" is an objective, measurable, and purely technical property. We challenge this view, showing that evaluating AI performance fundamentally depends on context-dependent normative decisions. These techno-normative choices are crucial for rigorous AI deployment, as they determine which errors are prioritised, how risks are distributed, and how trade-offs between competing objectives are resolved. This paper provides a legal-technical analysis of the choices that shape how accuracy is defined, measured, and assessed, using the 2024 European Union AI Act -- which mandates an "appropriate level of accuracy" for high-risk systems -- as a primary case study. We identify and analyse four choices central to any robust performance evaluation: (1) selecting metrics, (2) balancing multiple metrics, (3) measuring metrics against representative data, and (4) determining acceptance thresholds. For each choice, we study its relationship to the AI Act's accuracy requirement and associated documentation obligations, show how its technical implementation embeds implicit or explicit assumptions about acceptable risks, errors, and trade-offs, and discuss the implications for the practical implementation of the AI Act by examples and related technical standards. By making the techno-normative dimensions of accuracy explicit, this paper contributes to broader interdisciplinary debates on AI governance and regulation, and offers specific guidance for regulators, auditors, and developers tasked with translating (legal) safety requirements into technical practice.

Toward Artificial Intelligence Enabled Earth System Coupling

arXiv:2604.03289v1 Announce Type: cross Abstract: Coupling constitutes a foundational mechanism in the Earth system, regulating the interconnected physical, chemical, and biological processes that link its spheres. This review examines how emerging artificial intelligence (AI) methods create new opportunities to enhance Earth system coupling and address long-standing limitations in multi-component models. Rather than surveying next-generation modelling efforts broadly, we focus specifically on how state-of-the-art AI techniques can strengthen cross-domain interactions, support more coherent multi-component representations, and enable progress toward unified Earth system frameworks. The scope extends beyond climate models to include any modelling system in which Earth spheres interact. We outline emerging opportunities, persistent limitations, and conceptual pathways through which AI may enhance physical consistency, interpretability, and integration across domains. In doing so, this review provides a structured foundation for understanding the role of AI in advancing coupled Earth system modelling.

Towards Intelligent Energy Security: A Unified Spatio-Temporal and Graph Learning Framework for Scalable Electricity Theft Detection in Smart Grids

arXiv:2604.03344v1 Announce Type: cross Abstract: Electricity theft and non-technical losses (NTLs) remain critical challenges in modern smart grids, causing significant economic losses and compromising grid reliability. This study introduces the SmartGuard Energy Intelligence System (SGEIS), an integrated artificial intelligence framework for electricity theft detection and intelligent energy monitoring. The proposed system combines supervised machine learning, deep learning-based time-series modeling, Non-Intrusive Load Monitoring (NILM), and graph-based learning to capture both temporal and spatial consumption patterns. A comprehensive data processing pipeline is developed, incorporating feature engineering, multi-scale temporal analysis, and rule-based anomaly labeling. Deep learning models, including Long Short-Term Memory (LSTM), Temporal Convolutional Networks (TCN), and Autoencoders, are employed to detect abnormal usage patterns. In parallel, ensemble learning methods such as Random Forest, Gradient Boosting, XGBoost, and LightGBM are utilized for classification. To model grid topology and spatial dependencies, Graph Neural Networks (GNNs) are applied to identify correlated anomalies across interconnected nodes. The NILM module enhances interpretability by disaggregating appliance-level consumption from aggregate signals. Experimental results demonstrate strong performance, with Gradient Boosting achieving a ROC-AUC of 0.894, while graph-based models attain over 96% accuracy in identifying high-risk nodes. The hybrid framework improves detection robustness by integrating temporal, statistical, and spatial intelligence. Overall, SGEIS provides a scalable and practical solution for electricity theft detection, offering high accuracy, improved interpretability, and strong potential for real-world smart grid deployment.

Testing the Limits of Truth Directions in LLMs

arXiv:2604.03754v1 Announce Type: cross Abstract: Large language models (LLMs) have been shown to encode truth of statements in their activation space along a linear truth direction. Previous studies have argued that these directions are universal in certain aspects, while more recent work has questioned this conclusion drawing on limited generalization across some settings. In this work, we identify a number of limits of truth-direction universality that have not been previously understood. We first show that truth directions are highly layer-dependent, and that a full understanding of universality requires probing at many layers in the model. We then show that truth directions depend heavily on task type, emerging in earlier layers for factual and later layers for reasoning tasks; they also vary in performance across levels of task complexity. Finally, we show that model instructions dramatically affect truth directions; simple correctness evaluation instructions significantly affect the generalization ability of truth probes. Our findings indicate that universality claims for truth directions are more limited than previously known, with significant differences observable for various model layers, task difficulties, task types, and prompt templates.

CountsDiff: A Diffusion Model on the Natural Numbers for Generation and Imputation of Count-Based Data

arXiv:2604.03779v1 Announce Type: cross Abstract: Diffusion models have excelled at generative tasks for both continuous and token-based domains, but their application to discrete ordinal data remains underdeveloped. We present CountsDiff, a diffusion framework designed to natively model distributions on the natural numbers. CountsDiff extends the Blackout diffusion framework by simplifying its formulation through a direct parameterization in terms of a survival probability schedule and an explicit loss weighting. This introduces flexibility through design parameters with direct analogues in existing diffusion modeling frameworks. Beyond this reparameterization, CountsDiff introduces features from modern diffusion models, previously absent in counts-based domains, including continuous-time training, classifier-free guidance, and churn/remasking reverse dynamics that allow non-monotone reverse trajectories. We propose an initial instantiation of CountsDiff and validate it on natural image datasets (CIFAR-10, CelebA), exploring the effects of varying the introduced design parameters in a complex, well-studied, and interpretable data domain. We then highlight biological count assays as a natural use case, evaluating CountsDiff on single-cell RNA-seq imputation in a fetal cell and heart cell atlas. Remarkably, we find that even this simple instantiation matches or surpasses the performance of a state-of-the-art discrete generative model and leading RNA-seq imputation methods, while leaving substantial headroom for further gains through optimized design choices in future work.

What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features

arXiv:2604.04720v1 Announce Type: cross Abstract: Large Reasoning Models (LRMs) still exhibit large performance gaps between English and other languages, yet much current work assumes these gaps can be closed simply by making reasoning in every language resemble English reasoning. This work challenges this assumption by asking instead: what actually characterizes effective reasoning in multilingual settings, and to what extent do English-derived reasoning features genuinely help in other languages? We first define a suite of measurable reasoning features spanning multilingual alignment, reasoning step, and reasoning flow aspects of reasoning traces, and use logistic regression to quantify how each feature associates with final answer accuracy. We further train sparse autoencoders over multilingual traces to automatically discover latent reasoning concepts that instantiate or extend these features. Finally, we use the features as test-time selection policies to examine whether they can steer models toward stronger multilingual reasoning. Across two mathematical reasoning benchmarks, four LRMs, and 10 languages, we find that most features are positively associated with accuracy, but the strength of association varies considerably across languages and can even reverse in some. Our findings challenge English-centric reward designs and point toward adaptive objectives that accommodate language-specific reasoning patterns, with concrete implications for multilingual benchmark and reward design.

Individual and Combined Effects of English as a Second Language and Typos on LLM Performance

arXiv:2604.04723v1 Announce Type: cross Abstract: Large language models (LLMs) are used globally, and because much of their training data is in English, they typically perform best on English inputs. As a result, many non-native English speakers interact with them in English as a second language (ESL), and these inputs often contain typographical errors. Prior work has largely studied the effects of ESL variation and typographical errors separately, even though they often co-occur in real-world use. In this study, we use the Trans-EnV framework to transform standard English inputs into eight ESL variants and apply MulTypo to inject typos at three levels: low, moderate, and severe. We find that combining ESL variation and typos generally leads to larger performance drops than either factor alone, though the combined effect is not simply additive. This pattern is clearest on closed-ended tasks, where performance degradation can be characterized more consistently across ESL variants and typo levels, while results on open-ended tasks are more mixed. Overall, these findings suggest that evaluations on clean standard English may overestimate real-world model performance, and that evaluating ESL variation and typographical errors in isolation does not fully capture model behavior in realistic settings.

Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics

arXiv:2510.09901v2 Announce Type: replace Abstract: Computing has long served as a cornerstone of scientific discovery. Recently, a paradigm shift has emerged with the rise of large language models (LLMs), introducing autonomous systems, referred to as agents, that accelerate discovery across varying levels of autonomy. These language agents provide a flexible and versatile framework that orchestrates interactions with human scientists, natural language, computer language and code, and physics. This paper presents our view and vision of LLM-based scientific agents and their growing role in transforming the scientific discovery lifecycle, from hypothesis discovery, experimental design and execution, to result analysis and refinement. We critically examine current methodologies, emphasizing key innovations, practical achievements, and outstanding limitations. Additionally, we identify open research challenges and outline promising directions for building more robust, generalizable, and adaptive scientific agents. Our analysis highlights the transformative potential of autonomous agents to accelerate scientific discovery across diverse domains.

IV Co-Scientist: Multi-Agent LLM Framework for Causal Instrumental Variable Discovery

arXiv:2602.07943v2 Announce Type: replace Abstract: In the presence of confounding between an endogenous variable and the outcome, instrumental variables (IVs) are used to isolate the causal effect of the endogenous variable. Identifying valid instruments requires interdisciplinary knowledge, creativity, and contextual understanding, making it a non-trivial task. In this paper, we investigate whether large language models (LLMs) can aid in this task. We perform a two-stage evaluation framework. First, we test whether LLMs can recover well-established instruments from the literature, assessing their ability to replicate standard reasoning. Second, we evaluate whether LLMs can identify and avoid instruments that have been empirically or theoretically discredited. Building on these results, we introduce IV Co-Scientist, a multi-agent system that proposes, critiques, and refines IVs for a given treatment-outcome pair. We also introduce a statistical test to contextualize consistency in the absence of ground truth. Our results show the potential of LLMs to discover valid instrumental variables from a large observational database.

Isatuximab, carfilzomib, lenalidomide and dexamethasone in newly diagnosed multiple myeloma: a randomized phase 3 trial

Nature Medicine, Published online: 06 April 2026; doi:10.1038/s41591-026-04282-0

In the phase 3 EMN24 IsKia trial, transplant-eligible patients with newly diagnosed multiple myeloma who received isatuximab with carfilzomib, lenalidomide and dexamethasone pretransplant induction and post-transplant consolidation showed higher rates of measurable residual disease negativity after consolidation than patients who received carfilzomib, lenalidomide and dexamethasone.

Developing psychosocial phenotypes to understand engagement with digital health technologies for heart failure

npj Digital Medicine, Published online: 04 April 2026; doi:10.1038/s41746-026-02571-z

Developing psychosocial phenotypes to understand engagement with digital health technologies for heart failure

Translating ctDNA into cutaneous melanoma care: An international expert survey

Eur J Cancer. 2026 Mar 19;239:116676. doi: 10.1016/j.ejca.2026.116676. Online ahead of print.

ABSTRACT

BACKGROUND: Circulating tumor DNA (ctDNA) is a promising biomarker in melanoma, with higher sensitivity for tumor burden detection than conventional diagnostics. While well established in research, clinical routine implementation remains pending. Key global questions concern optimal clinical applications and barriers to adoption.

METHODS: A web-based survey of 116 members of the Melanoma World Society Study Group assessed international expert opinions on ctDNA utility across predefined clinical scenarios. The questionnaire included 18 general questions on ctDNA use and 5 clinical vignettes with de-identified patient data and retrospectively obtained ctDNA results.

RESULTS: ctDNA was rated most valuable for detecting minimal residual disease (mean score 3.63), surveillance of recurrent disease (3.85), and stage IV melanoma (3.82), with limited utility in early stages. Experts considered ctDNA superior to S100 and LDH for early relapse detection and identifying progressive disease. Most participants (80%) agreed that ctDNA correlates with radiographic response, and 82% favored its integration into routine follow-ups. In urgent high-tumor-burden settings, 82.8% would initiate BRAFi/MEKi therapy based on ctDNA if tissue analysis was pending, and 93.9% if unavailable. For central nervous system lesions, 62% did not support blood ctDNA, while 66% considered cerebrospinal fluid valuable. Pragmatic approaches with small to mid-size targeted panels and short turnaround times were preferred. Major barriers included the need for prospective trials (85%), standardized guidelines (83%), and reimbursement policies (82%).

CONCLUSION: Key opinion leaders regarded ctDNA as a valuable adjunct selected melanoma scenarios. Validation through prospective studies, guideline development, and reimbursement frameworks are essential for broader clinical implementation.

PMID:41932032 | DOI:10.1016/j.ejca.2026.116676

Integrating liquid biopsies and artificial intelligence for early cancer detection: A systematic review and meta-analysis

Eur J Cancer. 2026 Mar 24;239:116699. doi: 10.1016/j.ejca.2026.116699. Online ahead of print.

ABSTRACT

INTRODUCTION: The latest generation of liquid biopsies incorporates multi-omic features, including genomics, methylomics, and fragmentomics. Machine learning (ML) approaches have been proposed to synthesize these complex biological data for the development of diagnostic classifiers. This study aims to evaluate the integration of ML with circulating cell-free DNA (cfDNA) analysis for early cancer detection.

METHODS: Medline, Embase, Cochrane, and Web of Science were searched in July 2025. Eligible studies combined ML and cfDNA features to distinguish cancer patients (stages I-III) from non-cancer controls. Summary diagnostic performance metrics and their 95% confidence intervals (CI) were calculated.

RESULTS: The study included 109 articles permitting analyses for lung (n = 34), liver (n = 29), colorectal (n = 28), pancreatic (n = 16), breast (n = 17), esophageal (n = 12), ovarian (n = 13), gastric (n = 9), head and neck (n = 4), and mixed (n = 27) cancer types. Specificity was consistently high across all tumor types and stages (94%-99%). Sensitivity ranged from 72% to 92% for stage I-III, 44-91% for stage I, 71-98% for stage II and 83-99% for stage III. In the pooled study population, neural networks (90%, 95% CI: 81%-95%), random forest (86%, 95% CI: 77%-92%) and heterogeneous ensemble learning (85%, 95% CI: 79%-89%) demonstrated the highest sensitivity. The stratified analysis by classifier feature revealed 86% (95% CI: 80%-90%) sensitivity for fragmentation and 81% (95% CI: 76%-85%) for methylation, with 92%-96% specificity.

CONCLUSION: ML and cfDNA profiling show potential for early cancer detection, with ensemble methods, neural networks and random forests achieving the best overall performance. Fragmentomic features provide the highest sensitivity.

PMID:41930854 | DOI:10.1016/j.ejca.2026.116699

Integrating liquid biopsies and artificial intelligence for early cancer detection: A systematic review and meta-analysis

Eur J Cancer. 2026 Mar 24;239:116699. doi: 10.1016/j.ejca.2026.116699. Online ahead of print.

ABSTRACT

INTRODUCTION: The latest generation of liquid biopsies incorporates multi-omic features, including genomics, methylomics, and fragmentomics. Machine learning (ML) approaches have been proposed to synthesize these complex biological data for the development of diagnostic classifiers. This study aims to evaluate the integration of ML with circulating cell-free DNA (cfDNA) analysis for early cancer detection.

METHODS: Medline, Embase, Cochrane, and Web of Science were searched in July 2025. Eligible studies combined ML and cfDNA features to distinguish cancer patients (stages I-III) from non-cancer controls. Summary diagnostic performance metrics and their 95% confidence intervals (CI) were calculated.

RESULTS: The study included 109 articles permitting analyses for lung (n = 34), liver (n = 29), colorectal (n = 28), pancreatic (n = 16), breast (n = 17), esophageal (n = 12), ovarian (n = 13), gastric (n = 9), head and neck (n = 4), and mixed (n = 27) cancer types. Specificity was consistently high across all tumor types and stages (94%-99%). Sensitivity ranged from 72% to 92% for stage I-III, 44-91% for stage I, 71-98% for stage II and 83-99% for stage III. In the pooled study population, neural networks (90%, 95% CI: 81%-95%), random forest (86%, 95% CI: 77%-92%) and heterogeneous ensemble learning (85%, 95% CI: 79%-89%) demonstrated the highest sensitivity. The stratified analysis by classifier feature revealed 86% (95% CI: 80%-90%) sensitivity for fragmentation and 81% (95% CI: 76%-85%) for methylation, with 92%-96% specificity.

CONCLUSION: ML and cfDNA profiling show potential for early cancer detection, with ensemble methods, neural networks and random forests achieving the best overall performance. Fragmentomic features provide the highest sensitivity.

PMID:41930854 | DOI:10.1016/j.ejca.2026.116699

Integrating liquid biopsies and artificial intelligence for early cancer detection: A systematic review and meta-analysis

Eur J Cancer. 2026 Mar 24;239:116699. doi: 10.1016/j.ejca.2026.116699. Online ahead of print.

ABSTRACT

INTRODUCTION: The latest generation of liquid biopsies incorporates multi-omic features, including genomics, methylomics, and fragmentomics. Machine learning (ML) approaches have been proposed to synthesize these complex biological data for the development of diagnostic classifiers. This study aims to evaluate the integration of ML with circulating cell-free DNA (cfDNA) analysis for early cancer detection.

METHODS: Medline, Embase, Cochrane, and Web of Science were searched in July 2025. Eligible studies combined ML and cfDNA features to distinguish cancer patients (stages I-III) from non-cancer controls. Summary diagnostic performance metrics and their 95% confidence intervals (CI) were calculated.

RESULTS: The study included 109 articles permitting analyses for lung (n = 34), liver (n = 29), colorectal (n = 28), pancreatic (n = 16), breast (n = 17), esophageal (n = 12), ovarian (n = 13), gastric (n = 9), head and neck (n = 4), and mixed (n = 27) cancer types. Specificity was consistently high across all tumor types and stages (94%-99%). Sensitivity ranged from 72% to 92% for stage I-III, 44-91% for stage I, 71-98% for stage II and 83-99% for stage III. In the pooled study population, neural networks (90%, 95% CI: 81%-95%), random forest (86%, 95% CI: 77%-92%) and heterogeneous ensemble learning (85%, 95% CI: 79%-89%) demonstrated the highest sensitivity. The stratified analysis by classifier feature revealed 86% (95% CI: 80%-90%) sensitivity for fragmentation and 81% (95% CI: 76%-85%) for methylation, with 92%-96% specificity.

CONCLUSION: ML and cfDNA profiling show potential for early cancer detection, with ensemble methods, neural networks and random forests achieving the best overall performance. Fragmentomic features provide the highest sensitivity.

PMID:41930854 | DOI:10.1016/j.ejca.2026.116699

Integrating liquid biopsies and artificial intelligence for early cancer detection: A systematic review and meta-analysis

Eur J Cancer. 2026 Mar 24;239:116699. doi: 10.1016/j.ejca.2026.116699. Online ahead of print.

ABSTRACT

INTRODUCTION: The latest generation of liquid biopsies incorporates multi-omic features, including genomics, methylomics, and fragmentomics. Machine learning (ML) approaches have been proposed to synthesize these complex biological data for the development of diagnostic classifiers. This study aims to evaluate the integration of ML with circulating cell-free DNA (cfDNA) analysis for early cancer detection.

METHODS: Medline, Embase, Cochrane, and Web of Science were searched in July 2025. Eligible studies combined ML and cfDNA features to distinguish cancer patients (stages I-III) from non-cancer controls. Summary diagnostic performance metrics and their 95% confidence intervals (CI) were calculated.

RESULTS: The study included 109 articles permitting analyses for lung (n = 34), liver (n = 29), colorectal (n = 28), pancreatic (n = 16), breast (n = 17), esophageal (n = 12), ovarian (n = 13), gastric (n = 9), head and neck (n = 4), and mixed (n = 27) cancer types. Specificity was consistently high across all tumor types and stages (94%-99%). Sensitivity ranged from 72% to 92% for stage I-III, 44-91% for stage I, 71-98% for stage II and 83-99% for stage III. In the pooled study population, neural networks (90%, 95% CI: 81%-95%), random forest (86%, 95% CI: 77%-92%) and heterogeneous ensemble learning (85%, 95% CI: 79%-89%) demonstrated the highest sensitivity. The stratified analysis by classifier feature revealed 86% (95% CI: 80%-90%) sensitivity for fragmentation and 81% (95% CI: 76%-85%) for methylation, with 92%-96% specificity.

CONCLUSION: ML and cfDNA profiling show potential for early cancer detection, with ensemble methods, neural networks and random forests achieving the best overall performance. Fragmentomic features provide the highest sensitivity.

PMID:41930854 | DOI:10.1016/j.ejca.2026.116699

❌