❌

Reading view

Genomically matched therapy in advanced solid tumors: the randomized phase 2 ROME trial

Nature Medicine, Published online: 29 September 2025; doi:10.1038/s41591-025-03918-x

In the proof-of-concept phase 2 ROME trial, comprehensive genomic profiling followed by molecular tumor board evaluation and randomization of patients with metastatic solid cancer to receive personalized therapy or standard of care led to a significantly higher objective response rate and longer progression-free survival in patients who received personalized therapy.
  •  

Multi-Omics Feature Selection to Identify Biomarkers for Hepatocellular Carcinoma

Metabolites. 2025 Aug 28;15(9):575. doi: 10.3390/metabo15090575.

ABSTRACT

INTRODUCTION: Hepatocellular carcinoma (HCC), the most prevalent form of liver cancer, ranks as the third leading cause of mortality globally. Patients diagnosed with HCC exhibit a dismal prognosis mostly due to the emergence of symptoms in the advanced stages of the disease. Moreover, conventional biomarkers demonstrate insufficient efficacy in the early detection of HCC, hence highlighting the need for the identification of novel and more effective biomarkers.

METHODS: In this paper, we investigate methods for integration of multi-omics data we generated by both untargeted and targeted mass spectrometric analysis of serum samples from HCC cases and patients with liver cirrhosis. Specifically, the performances of several feature selection methods are evaluated on their abilities to identify a panel of multi-omics features that distinguish HCC cases from cirrhotic controls.

RESULTS: The integrative analysis identified key molecules associated with liver including such as leucine and isoleucine as well as SERPINA1, which is involved in LXR/RXR Activation and Acute Response signaling. A new method that uses recursive feature selection in conjunction with a transformer-based deep learning model as an estimator led to more promising results compared to other deep learning methods that perform disease classification and feature selection sequentially.

CONCLUSIONS: The findings in this study reinforce the importance of adapting or extending deep learning models to support robust feature selection, especially for integration of multi-omics data with limited sample size to avoid the risk of overfitting and the need for evaluation of the multi-omics features discovered in this study via blood samples from a larger and independent cohort to identify robust biomarkers for HCC.

PMID:41002959 | PMC:PMC12471784 | DOI:10.3390/metabo15090575

  •  

Association of HTR1F with Prognosis, Tumor Immune Microenvironment, and Drug Sensitivity in Cancer: A Multi-Omics Perspective

Biomedicines. 2025 Sep 11;13(9):2238. doi: 10.3390/biomedicines13092238.

ABSTRACT

Background:HTR1F (5-Hydroxytryptamine Receptor 1F) encodes a G protein-coupled receptor involved in serotonin signaling. Although dysregulated HTR1F expression has been implicated in certain malignancies, its biological functions and clinical significance across cancer types remain largely unexplored. Methods: We performed an integrative pan-cancer analysis of transcriptomic and pharmacogenomic datasets covering 34 cancer types (PAN-CAN cohort, N = 19,131; normal tissues, G = 60,499). Drug sensitivity and molecular docking analyses were conducted using the GSCALite database. The protein-protein interaction (PPI) network of HTR1F was constructed via the STRING database. Additionally, we evaluated the effects of HTR1F overexpression on proliferation and invasion in human lung squamous cell carcinoma (LUSC) cell lines NCI-H520 and NCI-H226. Results:HTR1F expression was significantly upregulated in 17 cancer types and was associated with poor prognosis, with LUSC showing an AUC of 0.912 for 1-year survival prediction. In LUSC, 695 genes were upregulated and 67 downregulated in response to HTR1F overexpression. HTR1F expression correlated with immune-related genes, immune checkpoints, tumor-infiltrating immune cells, tumor mutation burden (TMB), microsatellite instability (MSI), and drug responses. Genomic alterations, including amplification and deletion, were positively associated with HTR1F expression. Drug sensitivity analysis identified compounds such as sotrastaurin (-10.2 kcal/mol), austocystin D (-9.7 kcal/mol), and tivozanib (-9.3 kcal/mol) as potentially effective inhibitors based on predicted binding affinity. Functional enrichment analyses (GO, KEGG) and GSEA revealed that HTR1F is primarily involved in cell cycle regulation, DNA replication, cellular senescence, and immune-related pathways. Functional validation showed that HTR1F overexpression promotes proliferation of LUSC cells via the MAPK signaling pathway. Conclusions: Our integrative analysis highlights HTR1F as a potential biomarker associated with prognosis, immune modulation, and drug sensitivity across multiple cancer types. These findings provide a foundation for future experimental and clinical studies to explore HTR1F-targeted therapies.

PMID:41007799 | PMC:PMC12467612 | DOI:10.3390/biomedicines13092238

  •  

Using Large Language Models to Assess the Consistency of Randomized Controlled Trials on AI Interventions With CONSORT-AI: Cross-Sectional Survey

Background: Chatbots based on large language models (LLMs) have shown promise in evaluating the consistency of research. Previously, researchers used LLM to assess if randomized controlled trial (RCT) abstracts adhered to the CONSORT-Abstract guidelines. However, the consistency of artificial intelligence (AI) interventional RCTs align with the CONSORT-AI standards by LLMs remains unclear. Objective: The aim of this study is to identify the consistency of randomized controlled trials on AI interventions with CONSORT-AI using chatbots based on LLMs. Methods: This cross-sectional study employed six LLM models to assess the consistency of RCTs on AI interventions. The sample selection is based on articles published in JAMA Network Open, which included a total of 41 RCTs. All queries were submitted to LLMs through an API interface with a temperature setting of 0 to ensure deterministic responses. One researcher posed the questions to each model, while another independently verified the responses for validity before recording the results. The Overall Consistency Score (OCS), recall, inter-rater reliability and consistency of contents were analyzed. Results: We found gpt-4-0125-preview has the best average OCS (86.5%, 95%CI: 82.5%-90.5% and 81.6%, 95% CI: 77.6%-85.6%), followed by gpt-4-1106-preview(80.3%, 95%CI: 76.3%-84.3% and 78.0%, 95% CI: 74.0%-82.0%). The model with the worst average OCS is gpt-3.5-turbo-0125 (61.9%, 95%CI: 57.9%-65.9% and 63.0%, 95% CI: 59.0%-67.0%). Among the 11 unique items of CONSORT-AI, Item 2 (“State the inclusion and exclusion criteria at the level of the input data”) received the poorest overall evaluation across six models, with an average OCS of 48.8%. For other items, those with an average OCS greater than 80% across the six models included Items 1, 5, 8, and 9. Conclusions: GPT-4 variants demonstrate strong performance in assessing the consistency of RCTs with CONSORT-AI. Nonetheless, refining the prompts could enhance the precision and consistency of the outcomes. While AI tools like GPT-4 variants are valuable, they are not yet fully autonomous in addressing complex and nuanced tasks such as adherence to CONSORT-AI standards. Therefore, integrating AI with higher levels of human supervision and expertise will be crucial to ensuring more reliable and efficient evaluations, ultimately advancing the quality of medical research.
  •  

Integrative Spatial Omics for Systems-Level Mapping of Pathological Niches

bioRxiv [Preprint]. 2025 Sep 17:2025.09.12.675904. doi: 10.1101/2025.09.12.675904.

ABSTRACT

Spatial 'omics technologies are a powerful tool for mapping the relationship between cellular organization and molecular distributions in healthy and diseased tissue microenvironments. Here, we describe a novel multimodal pipeline that represents experimental and computational advances for spatiomolecular analysis of tissue samples across molecular classes. This adaptable method integrates matrix-assisted laser desorption/ionization (MALDI) imaging mass spectrometry (IMS) lipidomics, spatial transcriptomics (ST), multiplexed immunofluorescence microscopy (MxIF), and histopathological staining to uncover spatiomolecular profiles associated with unique cellular niches and pathological features. We demonstrate the power of this approach using two different complex human disease systems: Alzheimer's disease in human brain tissue and type 2 diabetes mellitus in the human pancreas. By identifying molecular markers associated with disease pathology in the pancreas and brain, we shed light on biologically significant pathways that are impacted in these two spatially complex diseases and highlight the powerful potential of accurate, high-resolution multimodal integration approaches.

PMID:41000710 | PMC:PMC12458195 | DOI:10.1101/2025.09.12.675904

  •  

FUSION: a web-based application for in-depth exploration of multi-omics data with brightfield histology

Nat Commun. 2025 Sep 25;16(1):8388. doi: 10.1038/s41467-025-63050-9.

ABSTRACT

Spatial technologies examining the cell and tissue microenvironment at near single-cell resolution are revealing important molecular insights. However, few tools enable integrated, interactive analysis of spatial-omics with tissue morphology in the same functional tissue unit. Here, we present FUSION (Functional Unit State Identification in Whole Slide Images), a web-based platform for visualizing and analyzing spatial-omics data with high-resolution histology. FUSION provides workflows for assessing cell compositions, quantitative morphometrics, and comparative tissue analyses. We demonstrate applicability across spatial assays, including 10x Visium, Visium HD, 10x Xenium, Cell DIVE, and PhenoCycler, applied to healthy and diseased tissues from kidney, small intestine, lung, and skin in the Human BioMolecular Atlas Program. FUSION is cloud-based, open-source, and accessible at https://fusion.hubmapconsortium.org/ , hosting over 50 paired datasets and tutorials. In a series of use cases, we show its capacity to distinguish renal glomeruli injury states, quantify morphometric changes, and characterize fibrosis with immune infiltration.

PMID:40998789 | PMC:PMC12462499 | DOI:10.1038/s41467-025-63050-9

  •  

Application of Nudges to Design Clinical Decision Support Tools: Systematic Approach Guided by Implementation Science

Background: Clinical decision support (CDS) is one strategy to increase evidence-based practices by clinicians. Despite its potential, CDS tools have had mixed results and are often disliked by clinicians. Principles from behavioral economics, including “nudges,” may improve the effectiveness and clinician satisfaction of CDS tools. Objective: This paper outlines a pragmatic approach grounded in implementation science to identify and prioritize how to incorporate different types of nudges into CDS tools. Methods: We applied the Messenger, Incentives, Norms, Defaults, Salience, Priming, Affect, Commitments and Ego (MINDSPACE) nudge framework and the Practical, Robust Implementation and Sustainability Model (PRISM) implementation science framework to systematically and pragmatically identify and prioritize different types of nudges for CDS tools. A case example of a CDS tool to improve guideline-concordant prescribing for patients with heart failure was used to illustrate how these frameworks can be applied in real-life scenarios. We describe a process of how these frameworks can be used pragmatically by clinicians and informaticists or more technical CDS builders to apply nudge theory to CDS tools. Results: Four iterative steps guided by PRISM were defined: 1) engage partners for user-centered design, 2) develop a shared understanding of the nudge types, 3) determine the nudge type for the overarching CDS format, and 4) brainstorm and prioritize nudge types and forms to address each modifiable contextual issue. These steps are iterative and intended to be adapted to align with the local resources and needs of various clinical scenarios and settings. We provide illustrative examples of how this approach was applied to the case example, including who we engaged, details of nudge design decisions, and lessons learned. Conclusions: We present a pragmatic approach to guide the selection and prioritization of nudges, informed by implementation science. This approach can be used to comprehensively and systematically consider key issues in designing CDS to optimize clinician satisfaction, effectiveness, equity, and sustainability while minimizing the potential for unintended consequences. The findings can be adapted and generalized to other health settings and clinical situations, advancing the goals of learning health systems to expedite the translation of evidence into practice.
  •  

Understanding the Role of Clinical Decision Support Systems Among Hospital Nurses Using the FITT (Fit Between Individuals, Tasks, and Technology) Framework: Qualitative Study

Background: Clinical decision support systems (CDSSs) have gained prominence in health care, aiding professionals in decision-making and improving patient outcomes. While physicians often use CDSSs for diagnosis and treatment optimization, nurses rely on these systems for tasks such as patient monitoring, prioritization, and care planning. In nursing practice, CDSSs can assist with timely detection of clinical deterioration, support infection control, and streamline care documentation. Despite their potential, the adoption and use of CDSSs by nurses face diverse challenges. Barriers such as alarm fatigue, limited usability, lack of integration with workflows, and insufficient training continue to undermine effective implementation. In contrast to the relatively extensive body of research on CDSS use by physicians, studies focusing on nurses remain limited, leaving a gap in understanding the unique facilitators and barriers they encounter. Objective: This study aimed to explore the facilitators and barriers influencing the adoption and use of CDSSs by nurses in hospitals, using an extended Fit Between Individuals, Tasks, and Technology (FITT) framework. Methods: A qualitative study was conducted using semistructured interviews with 22 nurses from across the Netherlands, representing 3 hospital types: general (n=9), top-clinical (n=12), and academic (n=1). The sample included a diverse mix of practicing nurses, nurses-in-training, and clinical nurse information officers, with clinical experience ranging from 1.5 to 38 years. Interview transcripts were analyzed thematically, beginning with an inductive coding approach to identify key factors. These were then categorized deductively using the extended FITT framework. In total, 988 code instances were examined. To ensure analytical rigor, the coding process was separately conducted by 2 researchers and reviewed by an expert panel. Results: A total of 26 distinct factors were identified, categorized into 4 FITT dimensions: technology-individual, technology-task, task-individual, and organizational context. Of these, 11 factors were facilitators (eg, cognition, clarification, and prevention), 7 were barriers (eg, alarm fatigue, poor design, and limited digital proficiency), and 8 were both facilitators and barriers depending on the context (eg, acceptance, workload, and training). In addition, key value tensions emerged, such as the balance between standardization and professional autonomy, and the trade-off between enhanced decision support and increased administrative burden. Conclusions: The findings underscore the complexity of CDSS adoption in nursing practice, highlighting the interaction of facilitators and barriers across FITT dimensions. Practical recommendations include participatory design processes, targeted training programs, advanced alert management systems, and strong organizational support. Addressing value tensions and aligning CDSS functionality with nurses’ workflows can enhance adoption and optimize patient outcomes. Trial Registration:
  •  

Multimodal foundation model and benchmark for comprehensive retinal OCT image analysis

npj Digital Medicine, Published online: 25 September 2025; doi:10.1038/s41746-025-01852-3

Multimodal foundation model and benchmark for comprehensive retinal OCT image analysis
  •  

Quality safety and disparity of an AI chatbot in managing chronic diseases: simulated patient experiments

npj Digital Medicine, Published online: 25 September 2025; doi:10.1038/s41746-025-01956-w

Quality safety and disparity of an AI chatbot in managing chronic diseases: simulated patient experiments
  •  

Ophthalmic drug discovery and development using artificial intelligence and digital health technologies

npj Digital Medicine, Published online: 25 September 2025; doi:10.1038/s41746-025-01954-y

Ophthalmic drug discovery and development using artificial intelligence and digital health technologies
  •  
  •  

Expanding care coordination in an integrated health system through causal machine learning

npj Digital Medicine, Published online: 24 September 2025; doi:10.1038/s41746-025-01925-3

Expanding care coordination in an integrated health system through causal machine learning
  •  

Diabetic Foot Ulcer Classification Models Using Artificial Intelligence and Machine Learning Techniques: Systematic Review

Background: Diabetes-related foot ulceration (DFU) is a common complication of diabetes, with a significant impact on survival, health care costs, and health-related quality of life. The prognosis of DFU varies widely among individuals. The International Working Group on the Diabetic Foot recently updated their guidelines on how to classify ulcers using “classical” classification and scoring systems. No system was recommended for individual prognostication, and the group considered that more detail in ulcer characterization was needed and that machine learning (ML)–based models may be the solution. Despite advances in the field, no assessment of available evidence was done. Objective: This study aimed to identify and collect available evidence assessing the ability of ML-based models to predict clinical outcomes in people with DFU. Methods: We searched the MEDLINE database (PubMed), Scopus, Web of Science, and IEEE Xplore for papers published up to July 2023. Studies were eligible if they were anterograde analytical studies that examined the prognostic abilities of ML models in predicting clinical outcomes in a population that included at least 80% of adults with DFU. The literature was screened independently by 2 investigators (MMS and DAR or EH in the first phase, and MMS and MAS in the second phase) for eligibility criteria and data extracted. The risk of bias was evaluated using the Quality In Prognosis Studies tool and the Prediction model Risk Of Bias Assessment Tool by 2 investigators (MMS and MAS) independently. A narrative synthesis was conducted. Results: We retrieved a total of 2412 references after removing duplicates, of which 167 were subjected to full-text screening. Two references were added from searching relevant studies’ lists of references. A total of 11 studies, comprising 13 papers, were included focusing on 3 outcomes: wound healing, lower extremity amputation, and mortality. Overall, 55 predictive models were created using mostly clinical characteristics, random forest as the developing method, and area under the receiver operating characteristic curve (AUROC) as a discrimination accuracy measure. AUROC varied from 0.56 to 0.94, with the majority of the models reporting an AUROC equal or superior to 0.8 but lacking 95% CIs. All studies were found to have a high risk of bias, mainly due to a lack of uniform variable definitions, outcome definitions and follow-up periods, insufficient sample sizes, and inadequate handling of missing data. Conclusions: We identified several ML-based models predicting clinical outcomes with good discriminatory ability in people with DFU. Due to the focus on development and internal validation of the models, the proposal of several models in each study without selecting the “best one,” and the use of nonexplainable techniques, the use of this type of model is clearly impaired. Future studies externally validating explainable models are needed so that ML models can become a reality in DFU care. Trial Registration: PROSPERO CRD42022308248; https://www.crd.york.ac.uk/PROSPERO/view/CRD42022308248
  •  

Fine-Tuning Methods for Large Language Models in Clinical Medicine by Supervised Fine-Tuning and Direct Preference Optimization: Comparative Evaluation

Background: Large language model (LLM) fine tuning is the process of adjusting out-of-the-box model weights using a dataset of interest. Fine tuning can be a powerful technique to improve model performance in fields like medicine, where data access is restricted and LLMs may have poor out-of-the-box performance. Objective: In this study we investigated the benefits of fine tuning with supervised fine tuning (SFT) and direct preference optimization (DPO) across a range of LLM applications for medicine Methods: We use Llama3 7B and Mistral 7B v2 to compare the performance of SFT and DPO across four datasets for common natural language tasks in medicine. The tasks evaluated were simple classification, clinical reasoning, summarization, and clinical triage. Results: Clinical Reasoning accuracy increased 8% and 7% with DPO over SFT for Llama3 (p value 0.003) and Mistral2 (p value 0.004) respectively. Summarization quality, graded on a five point Likert scale, increased 0.13 and 0.10 for Llama3 and Mistral2 (p values
  •  

Comparative Evaluation of a Medical Large Language Model in Answering Real-World Radiation Oncology Questions: Multicenter Observational Study

Background: Large language models (LLMs) hold promise for supporting clinical tasks, particularly in data-driven and technical disciplines such as radiation oncology. While prior evaluation studies have focused on examination-style settings for evaluating LLMs, their performance in real-life clinical scenarios remains unclear. In the future, LLMs might be used as general AI assistants to answer questions arising in clinical practice. It is unclear how well a modern LLM, locally executed within the infrastructure of a hospital, would answer such questions compared with clinical experts. Objective: This study aimed to assess the performance of a locally deployed, state-of-the-art medical LLM in answering real-world clinical questions in radiation oncology compared with clinical experts. The aim was to evaluate the overall quality of answers, as well as the potential harmfulness of the answers if used for clinical decision-making. Methods: Physicians from 10 departments of European hospitals collected questions arising in the clinical practice of radiation oncology. Fifty of these questions were answered by 3 senior radiation oncology experts with at least 10 years of work experience, as well as the LLM Llama3-OpenBioLLM-70B (Ankit Pal and Malaikannan Sankarasubbu). In a blinded review, physicians rated the overall answer quality on a 5-point Likert scale (quality), assessed whether an answer might be potentially harmful if used for clinical decision-making (harmfulness), and determined if responses were from an expert or the LLM (recognizability). Comparisons between clinical experts and LLMs were then made for quality, harmfulness, and recognizability. Results: There were no significant differences between the quality of the answers between LLM and clinical experts (mean scores of 3.38 vs 3.63; median 4.00, IQR 3.00-4.00 vs median 3.67, IQR 3.33-4.00; P=.26; Wilcoxon signed rank test). The answers were deemed potentially harmful in 13% of cases for the clinical experts compared with 16% of cases for the LLM (P=.63; Fisher exact test). Physicians correctly identified whether an answer was given by a clinical expert or an LLM in 78% and 72% of cases, respectively. Conclusions: A state-of-the-art medical LLM can answer real-life questions from the clinical practice of radiation oncology similarly well as clinical experts regarding overall quality and potential harmfulness. Such LLMs can already be deployed within the local hospital environment at an affordable cost. While LLMs may not yet be ready for clinical implementation as general AI assistants, the technology continues to improve at a rapid pace. Evaluation studies based on real-life situations are important to better understand the weaknesses and limitations of LLMs in clinical practice. Such studies are also crucial to define when the technology is ready for clinical implementation. Furthermore, education for health care professionals on generative AI is needed to ensure responsible clinical implementation of this transforming technology.
  •  
❌