❌

Reading view

Liquid biopsy in gastrointestinal oncology: clinical applications and translational integration of ctDNA, CTCs, and sEVs

Oncol Rev. 2025 Oct 20;19:1702932. doi: 10.3389/or.2025.1702932. eCollection 2025.

ABSTRACT

BACKGROUND AND AIMS: Liquid biopsy offers a minimally invasive tool to detect actionable mutations, monitor minimal residual disease (MRD), and guide therapy in gastrointestinal (GI) cancers. We critically review the clinical utility of circulating tumor DNA (ctDNA), circulating tumor cells (CTCs), and small extracellular vesicles (sEVs) across GI malignancies and propose a framework for their integration into clinical practice.

METHODS: We synthesized evidence from over 200 studies, including prospective trials and translational research, to assess diagnostic accuracy, prognostic value, and clinical actionability of each biomarker type in esophageal, gastric, colorectal, pancreatic, hepatocellular, and biliary cancers.

RESULTS: ctDNA has shown strong potential for MRD detection and treatment monitoring, particularly in colorectal and pancreatic cancer. CTCs offer insights into metastatic risk and therapeutic resistance, while sEVs provide molecular cargo relevant to immunomodulation and disease progression. Emerging microfluidics and AI-driven multi-omics approaches may overcome current limitations.

CONCLUSION: The integration of liquid biopsy technologies into GI oncology holds promise for early detection and precision therapy. We propose a five-phase clinical roadmap and outine the key research gaps that need to be addressed before widespread implementation in routine care.

PMID:41190015 | PMC:PMC12580207 | DOI:10.3389/or.2025.1702932

  •  

Combining International Standards to Develop Clinical Decision Support for Parent Smoking Cessation in Pediatrics

Smoking has severe health consequences, and secondhand smoke (SHS) exposure among children increases the risk of sudden infant death syndrome, chronic respiratory diseases, such as asthma, and lung cancer in adulthood. For many parents, pediatricians are the primary source of interaction with the healthcare system. Nevertheless, in pediatric settings, appropriate tobacco treatments are rarely, if ever, provided to parents who smoke. To best address tobacco use among parents, it is ideal to develop scalable solutions that are coordinated across health systems, community partners, and national services within pediatric settings. We describe our experience developing and implementing a parent tobacco treatment platform (PTTP) within a pediatric institution that leverages multiple international standards to support interoperability, with the overarching goal of providing a model for how such work can be approached. The clinical decision support (CDS) system includes clinician- and patient-facing components, connects parents to three different treatment options (nicotine replacement therapy, text-based counseling, and telephonic counseling), and incorporates three international standards (Fast Healthcare Interoperability Resources [FHIR], SMART on FHIR, and CDS Hooks). FHIR is used across all components. SMART on FHIR is limited to the clinician-facing tool, and CDS Hooks is used in the patient-facing portion. While healthcare interoperability standards supported a significant portion of the overall system, non-standard technologies and enhancements of existing standards were also required. Further, no connections with community partners could use existing interoperability standards. Over one year, the CDS was used in 194,946 visits, identified 7,847 parents who smoke, and connected 2,954 parents to 6,320 distinct treatment services, a significant improvement compared to prior efforts. Our project demonstrates that building CDS systems using international standards, such as SMART on FHIR, FHIR, and CDS Hooks, is possible, but challenges remain. Limits in the CDS Hooks standard to support common workflows and a lack of communication standards used by 3rd parties outside the healthcare system represent areas for future work. To support these requirements, additional EHR-specific records and communication mechanisms were required.
  •  

Key Features of Digital Phenotyping for Monitoring Mental Disorders: Systematic Review

Background: The COVID-19 pandemic has intensified mental health issues globally, highlighting the urgent need for remote mental health monitoring. Digital phenotyping using smart devices has emerged as a promising approach, but it remains unclear which features are essential for predicting depression and anxiety. Objective: This systematic review aimed to identify the types of features collected through smart packages—integrated systems combining smartphones with wearable devices such as Actiwatches, smartbands, and smartwatches—and to determine which features should be considered essential for mental health monitoring based on the type of device used. Methods: A systematic review was conducted. Searches were performed across Web of Science, PubMed, and Scopus on February 5, 2025. Inclusion criteria comprised quantitative studies involving adults (≥19 years) using smart devices to predict depression or anxiety based on passive data collection. Studies focusing solely on smartphones or qualitative designs were excluded. Risk of bias was assessed using the Mixed Methods Appraisal Tool and Quality Criteria Checklist. Data were synthesized descriptively, and the relative contribution of each feature was further assessed by calculating coverage (proportion of studies using a feature) and importance among used (proportion identifying it as important when used). These metrics were visualized in quadrant-based scatter plots to identify consistently important features across devices. Results: From 1,382 records, 22 studies across 11 countries were included. The overall synthesis identified a core feature package—accelerometer (ACC), steps, Heart Rate (HR), and sleep. Device-specific analyses revealed further nuances: In Actiwatch studies, ACC and activity were consistently important, but sleep features were rarely examined. In smartbands, HR, steps, sleep, and phone usage were essential, while Global Positioning System (GPS), Electrodermal Activity (EDA), and skin temperature (TEMP) showed high importance when used, suggesting opportunities for broader adoption. In smartwatch studies, sleep and HR emerged as core features, whereas steps and ACC were widely used but often not identified as important. Conclusions: This systematic review identified a core feature package comprising ACC, steps, HR, and sleep that consistently contributes to mood disorder prediction across devices. At the same time, device-specific differences were observed: Actiwatch studies mainly emphasized ACC and activity but underutilized sleep features; smartbands highlighted HR, steps, sleep, and phone usage, with EDA, TEMP, and GPS showing additional promise; and smartwatches most reliably leveraged sleep and HR, while steps and ACC were widely used yet less effective. These findings suggest that while a shared core set of features exists, optimizing digital phenotyping requires tailoring feature selection to the characteristics of each device type. To advance this field, improving data accessibility, particularly in smartwatch ecosystems, and adopting standardized reporting frameworks will be essential to enhance comparability, reproducibility, and future meta-analytic integration. Clinical Trial: Open Science Framework https://osf.io/nz7k8
  •  

Curated and harmonised transcriptomics datasets of interstitial lung diseases

Data Brief. 2025 Oct 14;63:112139. doi: 10.1016/j.dib.2025.112139. eCollection 2025 Dec.

ABSTRACT

This study provides manually curated and homogenised transcriptomics data of interstitial lung disease (ILD) patients retrieved from the NCBI Gene Expression Omnibus and European Nucleotide Archive repositories. The compendium includes 30 transcriptomics datasets generated with DNA microarrays and RNA sequencing (RNA-seq) technologies for a total of 1371 samples. All the datasets underwent metadata curation and harmonisation, data quality check, and preprocessing with standardised procedures. Furthermore, a robust data model was developed to standardise phenotypic data, thereby enhancing comparability across heterogeneous datasets. Gene expression data and lists of differentially expressed genes computed between ILD and healthy samples are provided. Among the ILDs included in this study, idiopathic pulmonary fibrosis (IPF) is the most represented worldwide. Co-expression networks of IPF and healthy samples were inferred, which are also included in this study. This study enhances the Findability, Accessibility, Interoperability, and Reusability (FAIR) of publicly available transcriptomic datasets related to ILDs. The resulting resource provides a integrated platform for the implementation and validation of systems biology and pharmacology approaches, facilitating the development of novel diagnostic and therapeutic strategies for ILDs.

PMID:41189603 | PMC:PMC12581653 | DOI:10.1016/j.dib.2025.112139

  •  

A multimodal whole-slide foundation model for pathology

Nature Medicine, Published online: 05 November 2025; doi:10.1038/s41591-025-03982-3

Pretrained using 335,645 whole-slide images, a foundation model is developed to provide representations for slide- and patient-level tasks. It is capable of performing clinical tasks and generating reports even in data-scarce scenarios, such as rare cancer diagnosis and survival prediction, without requiring further fine-tuning.
  •  

Fair human-centric image dataset for ethical AI benchmarking

Nature, Published online: 05 November 2025; doi:10.1038/s41586-025-09716-2

The Fair Human-Centric Image Benchmark (FHIBE, pronounced ‘Feebee’)—an image dataset that implements best practices for consent, privacy, compensation, safety, diversity and utility—can be used responsibly as a fairness evaluation dataset for many human-centric computer vision applications.
  •  

Opinion: Is it ever OK for doctors to ‘fake’ CPR?

On TV, CPR looks like a miracle: a few light pushes on the chest, a couple of assisted breaths, and the person sputters back to life.

“CPR has been represented in the media and TV shows and all of these other places as a relatively innocuous intervention with high rates of success from which people recover with little problem,” professor Jason Adam Wasserman said on this episode of the “First Opinion Podcast.” In fact, it can be physically damaging — broken ribs, punctured lungs — and painful. And for patients who are already medically frail, it often fails.

Read the rest…

  •  

Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Network

arXiv:2511.02238v1 Announce Type: new Abstract: Novel research ideas play a critical role in advancing scientific inquiries. Recent advancements in Large Language Models (LLMs) have demonstrated their potential to generate novel research ideas by leveraging large-scale scientific literature. However, previous work in research ideation has primarily relied on simplistic methods, such as keyword co-occurrence or semantic similarity. These approaches focus on identifying statistical associations in the literature but overlook the complex, contextual relationships between scientific concepts, which are essential to effectively leverage knowledge embedded in human literature. For instance, papers that simultaneously mention "keyword A" and "keyword B" often present research ideas that integrate both concepts. Additionally, some LLM-driven methods propose and refine research ideas using the model's internal knowledge, but they fail to effectively utilize the scientific concept network, limiting the grounding of ideas in established research. To address these challenges, we propose the Deep Ideation framework to address these challenges, integrating a scientific network that captures keyword co-occurrence and contextual relationships, enriching LLM-driven ideation. The framework introduces an explore-expand-evolve workflow to iteratively refine research ideas, using an Idea Stack to track progress. A critic engine, trained on real-world reviewer feedback, guides the process by providing continuous feedback on the novelty and feasibility of ideas. Our experiments show that our approach improves the quality of generated ideas by 10.67% compared to other methods, with ideas surpassing top conference acceptance levels. Human evaluation highlights their practical value in scientific research, and ablation studies confirm the effectiveness of each component in the workflow. Code repo is available at https://github.com/kyZhao-1/Deep-Ideation.
  •  

Q-Sat AI: Machine Learning-Based Decision Support for Data Saturation in Qualitative Studies

arXiv:2511.01935v1 Announce Type: cross Abstract: The determination of sample size in qualitative research has traditionally relied on the subjective and often ambiguous principle of data saturation, which can lead to inconsistencies and threaten methodological rigor. This study introduces a new, systematic model based on machine learning (ML) to make this process more objective. Utilizing a dataset derived from five fundamental qualitative research approaches - namely, Case Study, Grounded Theory, Phenomenology, Narrative Research, and Ethnographic Research - we developed an ensemble learning model. Ten critical parameters, including research scope, information power, and researcher competence, were evaluated using an ordinal scale and used as input features. After thorough preprocessing and outlier removal, multiple ML algorithms were trained and compared. The K-Nearest Neighbors (KNN), Gradient Boosting (GB), Random Forest (RF), XGBoost, and Decision Tree (DT) algorithms showed the highest explanatory power (Test R2 ~ 0.85), effectively modeling the complex, non-linear relationships involved in qualitative sampling decisions. Feature importance analysis confirmed the vital roles of research design type and information power, providing quantitative validation of key theoretical assumptions in qualitative methodology. The study concludes by proposing a conceptual framework for a web-based computational application designed to serve as a decision support system for qualitative researchers, journal reviewers, and thesis advisors. This model represents a significant step toward standardizing sample size justification, enhancing transparency, and strengthening the epistemological foundation of qualitative inquiry through evidence-based, systematic decision-making.
  •  

Causal Graph Neural Networks for Healthcare

arXiv:2511.02531v1 Announce Type: cross Abstract: Healthcare artificial intelligence systems routinely fail when deployed across institutions, with documented performance drops and perpetuation of discriminatory patterns embedded in historical data. This brittleness stems, in part, from learning statistical associations rather than causal mechanisms. Causal graph neural networks address this triple crisis of distribution shift, discrimination, and inscrutability by combining graph-based representations of biomedical data with causal inference principles to learn invariant mechanisms rather than spurious correlations. This Review examines methodological foundations spanning structural causal models, disentangled causal representation learning, and techniques for interventional prediction and counterfactual reasoning on graphs. We analyse applications demonstrating clinical value across psychiatric diagnosis through brain network analysis, cancer subtyping via multi-omics causal integration, continuous physiological monitoring with mechanistic interpretation, and drug recommendation correcting prescription bias. These advances establish foundations for patient-specific Causal Digital Twins, enabling in silico clinical experimentation, with integration of large language models for hypothesis generation and causal graph neural networks for mechanistic validation. Substantial barriers remain, including computational requirements precluding real-time deployment, validation challenges demanding multi-modal evidence triangulation beyond cross-validation, and risks of causal-washing where methods employ causal terminology without rigorous evidentiary support. We propose tiered frameworks distinguishing causally-inspired architectures from causally-validated discoveries and identify critical research priorities making causal rather than purely associational claims.
  •  

AI Diffusion in Low Resource Language Countries

arXiv:2511.02752v1 Announce Type: cross Abstract: Artificial intelligence (AI) is diffusing globally at unprecedented speed, but adoption remains uneven. Frontier Large Language Models (LLMs) are known to perform poorly on low-resource languages due to data scarcity. We hypothesize that this performance deficit reduces the utility of AI, thereby slowing adoption in Low-Resource Language Countries (LRLCs). To test this, we use a weighted regression model to isolate the language effect from socioeconomic and demographic factors, finding that LRLCs have a share of AI users that is approximately 20% lower relative to their baseline. These results indicate that linguistic accessibility is a significant, independent barrier to equitable AI diffusion.
  •  

TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models

arXiv:2511.02802v1 Announce Type: cross Abstract: Tabular foundation models represent a growing paradigm in structured data learning, extending the benefits of large-scale pretraining to tabular domains. However, their adoption remains limited due to heterogeneous preprocessing pipelines, fragmented APIs, inconsistent fine-tuning procedures, and the absence of standardized evaluation for deployment-oriented metrics such as calibration and fairness. We present TabTune, a unified library that standardizes the complete workflow for tabular foundation models through a single interface. TabTune provides consistent access to seven state-of-the-art models supporting multiple adaptation strategies, including zero-shot inference, meta-learning, supervised fine-tuning (SFT), and parameter-efficient fine-tuning (PEFT). The framework automates model-aware preprocessing, manages architectural heterogeneity internally, and integrates evaluation modules for performance, calibration, and fairness. Designed for extensibility and reproducibility, TabTune enables consistent benchmarking of adaptation strategies of tabular foundation models. The library is open source and available at https://github.com/Lexsi-Labs/TabTune .
  •  

MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning

arXiv:2511.02805v1 Announce Type: cross Abstract: Typical search agents concatenate the entire interaction history into the LLM context, preserving information integrity but producing long, noisy contexts, resulting in high computation and memory costs. In contrast, using only the current turn avoids this overhead but discards essential information. This trade-off limits the scalability of search agents. To address this challenge, we propose MemSearcher, an agent workflow that iteratively maintains a compact memory and combines the current turn with it. At each turn, MemSearcher fuses the user's question with the memory to generate reasoning traces, perform search actions, and update memory to retain only information essential for solving the task. This design stabilizes context length across multi-turn interactions, improving efficiency without sacrificing accuracy. To optimize this workflow, we introduce multi-context GRPO, an end-to-end RL framework that jointly optimize reasoning, search strategies, and memory management of MemSearcher Agents. Specifically, multi-context GRPO samples groups of trajectories under different contexts and propagates trajectory-level advantages across all conversations within them. Trained on the same dataset as Search-R1, MemSearcher achieves significant improvements over strong baselines on seven public benchmarks: +11% on Qwen2.5-3B-Instruct and +12% on Qwen2.5-7B-Instruct relative average gains. Notably, the 3B-based MemSearcher even outperforms 7B-based baselines, demonstrating that striking a balance between information integrity and efficiency yields both higher accuracy and lower computational overhead. The code and models will be publicly available at https://github.com/icip-cas/MemSearcher
  •  

How can we assess human-agent interactions? Case studies in software agent design

arXiv:2510.09801v2 Announce Type: replace Abstract: LLM-powered agents are both a promising new technology and a source of complexity, where choices about models, tools, and prompting can affect their usefulness. While numerous benchmarks measure agent accuracy across domains, they mostly assume full automation, failing to represent the collaborative nature of real-world use cases. In this paper, we make two major steps towards the rigorous assessment of human-agent interactions. First, we propose PULSE, a framework for more efficient human-centric evaluation of agent designs, which comprises collecting user feedback, training an ML model to predict user satisfaction, and computing results by combining human satisfaction ratings with model-generated pseudo-labels. Second, we deploy the framework on a large-scale web platform built around the open-source software agent OpenHands, collecting in-the-wild usage data across over 15k users. We conduct case studies around how three agent design decisions -- choice of LLM backbone, planning strategy, and memory mechanisms -- impact developer satisfaction rates, yielding practical insights for software agent design. We also show how our framework can lead to more robust conclusions about agent design, reducing confidence intervals by 40% compared to a standard A/B test. Finally, we find substantial discrepancies between in-the-wild results and benchmark performance (e.g., the anti-correlation between results comparing claude-sonnet-4 and gpt-5), underscoring the limitations of benchmark-driven evaluation. Our findings provide guidance for evaluations of LLM agents with humans and identify opportunities for better agent designs.
  •  

AutoPDL: Automatic Prompt Optimization for LLM Agents

arXiv:2504.04365v5 Announce Type: replace-cross Abstract: The performance of large language models (LLMs) depends on how they are prompted, with choices spanning both the high-level prompting pattern (e.g., Zero-Shot, CoT, ReAct, ReWOO) and the specific prompt content (instructions and few-shot demonstrations). Manually tuning this combination is tedious, error-prone, and specific to a given LLM and task. Therefore, this paper proposes AutoPDL, an automated approach to discovering good LLM agent configurations. Our approach frames this as a structured AutoML problem over a combinatorial space of agentic and non-agentic prompting patterns and demonstrations, using successive halving to efficiently navigate this space. We introduce a library implementing common prompting patterns using the PDL prompt programming language. AutoPDL solutions are human-readable, editable, and executable PDL programs that use this library. This approach also enables source-to-source optimization, allowing human-in-the-loop refinement and reuse. Evaluations across three tasks and seven LLMs (ranging from 3B to 70B parameters) show consistent accuracy gains ($9.21\pm15.46$ percentage points), up to 67.5pp, and reveal that selected prompting strategies vary across models and tasks.
  •  

From Passive to Proactive: A Multi-Agent System with Dynamic Task Orchestration for Intelligent Medical Pre-Consultation

arXiv:2511.01445v1 Announce Type: new Abstract: Global healthcare systems face critical challenges from increasing patient volumes and limited consultation times, with primary care visits averaging under 5 minutes in many countries. While pre-consultation processes encompassing triage and structured history-taking offer potential solutions, they remain limited by passive interaction paradigms and context management challenges in existing AI systems. This study introduces a hierarchical multi-agent framework that transforms passive medical AI systems into proactive inquiry agents through autonomous task orchestration. We developed an eight-agent architecture with centralized control mechanisms that decomposes pre-consultation into four primary tasks: Triage ($T_1$), History of Present Illness collection ($T_2$), Past History collection ($T_3$), and Chief Complaint generation ($T_4$), with $T_1$--$T_3$ further divided into 13 domain-specific subtasks. Evaluated on 1,372 validated electronic health records from a Chinese medical platform across multiple foundation models (GPT-OSS 20B, Qwen3-8B, Phi4-14B), the framework achieved 87.0% accuracy for primary department triage and 80.5% for secondary department classification, with task completion rates reaching 98.2% using agent-driven scheduling versus 93.1% with sequential processing. Clinical quality scores from 18 physicians averaged 4.56 for Chief Complaints, 4.48 for History of Present Illness, and 4.69 for Past History on a 5-point scale, with consultations completed within 12.7 rounds for $T_2$ and 16.9 rounds for $T_3$. The model-agnostic architecture maintained high performance across different foundation models while preserving data privacy through local deployment, demonstrating the potential for autonomous AI systems to enhance pre-consultation efficiency and quality in clinical settings.
  •  

Digital Twin based Automatic Reconfiguration of Robotic Systems in Smart Environments

arXiv:2511.00094v1 Announce Type: cross Abstract: Robotic systems have become integral to smart environments, enabling applications ranging from urban surveillance and automated agriculture to industrial automation. However, their effective operation in dynamic settings - such as smart cities and precision farming - is challenged by continuously evolving topographies and environmental conditions. Traditional control systems often struggle to adapt quickly, leading to inefficiencies or operational failures. To address this limitation, we propose a novel framework for autonomous and dynamic reconfiguration of robotic controllers using Digital Twin technology. Our approach leverages a virtual replica of the robot's operational environment to simulate and optimize movement trajectories in response to real-world changes. By recalculating paths and control parameters in the Digital Twin and deploying the updated code to the physical robot, our method ensures rapid and reliable adaptation without manual intervention. This work advances the integration of Digital Twins in robotics, offering a scalable solution for enhancing autonomy in smart, dynamic environments.
  •  

Diffusion Models at the Drug Discovery Frontier: A Review on Generating Small Molecules versus Therapeutic Peptides

arXiv:2511.00209v1 Announce Type: cross Abstract: Diffusion models have emerged as a leading framework in generative modeling, showing significant potential to accelerate and transform the traditionally slow and costly process of drug discovery. This review provides a systematic comparison of their application in designing two principal therapeutic modalities: small molecules and therapeutic peptides. We analyze how a unified framework of iterative denoising is adapted to the distinct molecular representations, chemical spaces, and design objectives of each modality. For small molecules, these models excel at structure-based design, generating novel, pocket-fitting ligands with desired physicochemical properties, yet face the critical hurdle of ensuring chemical synthesizability. Conversely, for therapeutic peptides, the focus shifts to generating functional sequences and designing de novo structures, where the primary challenges are achieving biological stability against proteolysis, ensuring proper folding, and minimizing immunogenicity. Despite these distinct challenges, both domains face shared hurdles: the need for more accurate scoring functions, the scarcity of high-quality experimental data, and the crucial requirement for experimental validation. We conclude that the full potential of diffusion models will be unlocked by bridging these modality-specific gaps and integrating them into automated, closed-loop Design-Build-Test-Learn (DBTL) platforms, thereby shifting the paradigm from chemical exploration to the targeted creation of novel therapeutics.
  •  

Investigating Label Bias and Representational Sources of Age-Related Disparities in Medical Segmentation

arXiv:2511.00477v1 Announce Type: cross Abstract: Algorithmic bias in medical imaging can perpetuate health disparities, yet its causes remain poorly understood in segmentation tasks. While fairness has been extensively studied in classification, segmentation remains underexplored despite its clinical importance. In breast cancer segmentation, models exhibit significant performance disparities against younger patients, commonly attributed to physiological differences in breast density. We audit the MAMA-MIA dataset, establishing a quantitative baseline of age-related bias in its automated labels, and reveal a critical Biased Ruler effect where systematically flawed labels for validation misrepresent a model's actual bias. However, whether this bias originates from lower-quality annotations (label bias) or from fundamentally more challenging image characteristics remains unclear. Through controlled experiments, we systematically refute hypotheses that the bias stems from label quality sensitivity or quantitative case difficulty imbalance. Balancing training data by difficulty fails to mitigate the disparity, revealing that younger patient cases are intrinsically harder to learn. We provide direct evidence that systemic bias is learned and amplified when training on biased, machine-generated labels, a critical finding for automated annotation pipelines. This work introduces a systematic framework for diagnosing algorithmic bias in medical segmentation and demonstrates that achieving fairness requires addressing qualitative distributional differences rather than merely balancing case counts.
  •  
❌