❌

Normal view

A Field Guide to Deploying AI Agents in Clinical Practice

arXiv:2509.26153v3 Announce Type: replace Abstract: Large language models (LLMs) integrated into agent-driven workflows hold immense promise for healthcare, yet a significant gap exists between their potential and practical implementation within clinical settings. To address this, we present a practitioner-oriented field manual for deploying generative agents that use electronic health record (EHR) data. This guide is informed by our experience deploying the "irAE-Agent", an automated system to detect immune-related adverse events from clinical notes at Mass General Brigham, and by structured interviews with 21 clinicians, engineers, and informatics leaders involved in the project. Our analysis reveals a critical misalignment in clinical AI development: less than 20% of our effort was dedicated to prompt engineering and model development, while over 80% was consumed by the sociotechnical work of implementation. We distill this effort into five "heavy lifts": data integration, model validation, ensuring economic value, managing system drift, and governance. By providing actionable solutions for each of these challenges, this field manual shifts the focus from algorithmic development to the essential infrastructure and implementation work required to bridge the "valley of death" and successfully translate generative AI from pilot projects into routine clinical care.

MARIA: A Framework for Marginal Risk Assessment without Ground Truth in AI Systems

arXiv:2510.27163v1 Announce Type: cross Abstract: Before deploying an AI system to replace an existing process, it must be compared with the incumbent to ensure improvement without added risk. Traditional evaluation relies on ground truth for both systems, but this is often unavailable due to delayed or unknowable outcomes, high costs, or incomplete data, especially for long-standing systems deemed safe by convention. The more practical solution is not to compute absolute risk but the difference between systems. We therefore propose a marginal risk assessment framework, that avoids dependence on ground truth or absolute risk. It emphasizes three kinds of relative evaluation methodology, including predictability, capability and interaction dominance. By shifting focus from absolute to relative evaluation, our approach equips software teams with actionable guidance: identifying where AI enhances outcomes, where it introduces new risks, and how to adopt such systems responsibly.

When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior

npj Digital Medicine, Published online: 17 October 2025; doi:10.1038/s41746-025-02008-z

When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior

Integrating multiomics analysis and machine learning to refine the molecular subtyping and prognostic analysis of stomach adenocarcinoma

Sci Rep. 2025 Jan 30;15(1):3843. doi: 10.1038/s41598-025-87444-3.

ABSTRACT

Stomach adenocarcinoma (STAD) is a common malignancy with high heterogeneity and a lack of highly precise treatment options. We downloaded the multiomics data of STAD patients in The Cancer Genome Atlas (TCGA)-STAD cohort, which included mRNA, microRNA, long non-coding RNA, somatic mutation, and DNA methylation data, from the sxdyc website. We synthesized the multiomics data of patients with STAD using 10 clustering methods, construct a consensus machine learning-driven signature (CMLS)-related prognostic models by combining 10 machine learning methods, and evaluated the prognosis models using the C-index. The prognostic relationship between CMLS and STAD was assessed using Kaplan-Meier curves, and the independent prognostic value of CMLS was determined by univariate and multivariate regression analyses. we also evaluated the immune characteristics, immunotherapy response, and drug sensitivity of different CMLS groups. The results of the multiomics analysis classified STAD into three subtypes, with CS1 resulting in the best survival outcome. In total, 10 hub genes (CES3, AHCYL2, APOD, EFEMP1, CYP1B1, ASPN, CPE, CLIP3, MAP1B, and DKK1) were screened and constructed the CMLS was significantly correlated with prognosis in patients with STAD and was an independent prognostic factor for patients with STAD. Using the CMLS risk score, all patients were divided into a high CMLS group and a low CMLS group. Patients in the low-CMLS group had better survival, more enriched immune cells, and higher tumor mutation load scores, suggesting better immunotherapy responsiveness and a possible "hot tumor" phenotype. Patients in the high-CMLS group had a significantly poorer prognosis and were less sensitive to immunotherapy but were likely to benefit more from chemotherapy and targeted therapy. In this study, 10 clustering methods and 10 machine learning methods were combined to analyze the multiomics of STAD, classify STAD into three subtypes, and constructed CMLS-related prognostic model features, which are important for accurate management and effective treatment of STAD.

PMID:39885324 | DOI:10.1038/s41598-025-87444-3

❌