❌

Normal view

Toward an AI Reasoning-Enabled System for Patient-Clinical Trial Matching

arXiv:2512.08026v1 Announce Type: new Abstract: Screening patients for clinical trial eligibility remains a manual, time-consuming, and resource-intensive process. We present a secure, scalable proof-of-concept system for Artificial Intelligence (AI)-augmented patient-trial matching that addresses key implementation challenges: integrating heterogeneous electronic health record (EHR) data, facilitating expert review, and maintaining rigorous security standards. Leveraging open-source, reasoning-enabled large language models (LLMs), the system moves beyond binary classification to generate structured eligibility assessments with interpretable reasoning chains that support human-in-the-loop review. This decision support tool represents eligibility as a dynamic state rather than a fixed determination, identifying matches when available and offering actionable recommendations that could render a patient eligible in the future. The system aims to reduce coordinator burden, intelligently broaden the set of trials considered for each patient and guarantee comprehensive auditability of all AI-generated outputs.

Principles2Plan: LLM-Guided System for Operationalising Ethical Principles into Plans

arXiv:2512.08536v1 Announce Type: new Abstract: Ethical awareness is critical for robots operating in human environments, yet existing automated planning tools provide little support. Manually specifying ethical rules is labour-intensive and highly context-specific. We present Principles2Plan, an interactive research prototype demonstrating how a human and a Large Language Model (LLM) can collaborate to produce context-sensitive ethical rules and guide automated planning. A domain expert provides the planning domain, problem details, and relevant high-level principles such as beneficence and privacy. The system generates operationalisable ethical rules consistent with these principles, which the user can review, prioritise, and supply to a planner to produce ethically-informed plans. To our knowledge, no prior system supports users in generating principle-grounded rules for classical planning contexts. Principles2Plan showcases the potential of human-LLM collaboration for making ethical automated planning more practical and feasible.

Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture

arXiv:2512.08130v1 Announce Type: cross Abstract: Both model developers and policymakers seek to quantify and mitigate the risk of rapidly-evolving frontier artificial intelligence (AI) models, especially large language models (LLMs), to facilitate bioterrorism or access to biological weapons. An important element of such efforts is the development of model benchmarks that can assess the biosecurity risk posed by a particular model. This paper describes the first component of a novel Biothreat Benchmark Generation (BBG) Framework. The BBG approach is designed to help model developers and evaluators reliably measure and assess the biosecurity risk uplift and general harm potential of existing and future AI models, while accounting for key aspects of the threat itself that are often overlooked in other benchmarking efforts, including different actor capability levels, and operational (in addition to purely technical) risk factors. As a pilot, the BBG is first being developed to address bacterial biological threats only. The BBG is built upon a hierarchical structure of biothreat categories, elements and tasks, which then serves as the basis for the development of task-aligned queries. This paper outlines the development of this biothreat task-query architecture, which we have named the Bacterial Biothreat Schema, while future papers will describe follow-on efforts to turn queries into model prompts, as well as how the resulting benchmarks can be implemented for model evaluation. Overall, the BBG Framework, including the Bacterial Biothreat Schema, seeks to offer a robust, re-usable structure for evaluating bacterial biological risks arising from LLMs across multiple levels of aggregation, which captures the full scope of technical and operational requirements for biological adversaries, and which accounts for a wide spectrum of biological adversary capabilities.

A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties

arXiv:2512.08185v1 Announce Type: cross Abstract: Medical Large Language Models (LLMs) are increasingly deployed for clinical decision support across diverse specialties, yet systematic evaluation of their robustness to adversarial misuse and privacy leakage remains inaccessible to most researchers. Existing security benchmarks require GPU clusters, commercial API access, or protected health data -- barriers that limit community participation in this critical research area. We propose a practical, fully reproducible framework for evaluating medical AI security under realistic resource constraints. Our framework design covers multiple medical specialties stratified by clinical risk -- from high-risk domains such as emergency medicine and psychiatry to general practice -- addressing jailbreaking attacks (role-playing, authority impersonation, multi-turn manipulation) and privacy extraction attacks. All evaluation utilizes synthetic patient records requiring no IRB approval. The framework is designed to run entirely on consumer CPU hardware using freely available models, eliminating cost barriers. We present the framework specification including threat models, data generation methodology, evaluation protocols, and scoring rubrics. This proposal establishes a foundation for comparative security assessment of medical-specialist models and defense mechanisms, advancing the broader goal of ensuring safe and trustworthy medical AI systems.

ClinicalTrialsHub: Bridging Registries and Literature for Comprehensive Clinical Trial Access

arXiv:2512.08193v1 Announce Type: cross Abstract: We present ClinicalTrialsHub, an interactive search-focused platform that consolidates all data from ClinicalTrials.gov and augments it by automatically extracting and structuring trial-relevant information from PubMed research articles. Our system effectively increases access to structured clinical trial data by 83.8% compared to relying on ClinicalTrials.gov alone, with potential to make access easier for patients, clinicians, researchers, and policymakers, advancing evidence-based medicine. ClinicalTrialsHub uses large language models such as GPT-5.1 and Gemini-3-Pro to enhance accessibility. The platform automatically parses full-text research articles to extract structured trial information, translates user queries into structured database searches, and provides an attributed question-answering system that generates evidence-grounded answers linked to specific source sentences. We demonstrate its utility through a user study involving clinicians, clinical researchers, and PhD students of pharmaceutical sciences and nursing, and a systematic automatic evaluation of its information extraction and question answering capabilities.

Are generative AI text annotations systematically biased?

arXiv:2512.08404v1 Announce Type: cross Abstract: This paper investigates bias in GLLM annotations by conceptually replicating manual annotations of Boukes (2024). Using various GLLMs (Llama3.1:8b, Llama3.3:70b, GPT4o, Qwen2.5:72b) in combination with five different prompts for five concepts (political content, interactivity, rationality, incivility, and ideology). We find GLLMs perform adequate in terms of F1 scores, but differ from manual annotations in terms of prevalence, yield substantively different downstream results, and display systematic bias in that they overlap more with each other than with manual annotations. Differences in F1 scores fail to account for the degree of bias.

Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models III: Implementing the Bacterial Biothreat Benchmark (B3) Dataset

arXiv:2512.08459v1 Announce Type: cross Abstract: The potential for rapidly-evolving frontier artificial intelligence (AI) models, especially large language models (LLMs), to facilitate bioterrorism or access to biological weapons has generated significant policy, academic, and public concern. Both model developers and policymakers seek to quantify and mitigate any risk, with an important element of such efforts being the development of model benchmarks that can assess the biosecurity risk posed by a particular model. This paper discusses the pilot implementation of the Bacterial Biothreat Benchmark (B3) dataset. It is the third in a series of three papers describing an overall Biothreat Benchmark Generation (BBG) framework, with previous papers detailing the development of the B3 dataset. The pilot involved running the benchmarks through a sample frontier AI model, followed by human evaluation of model responses, and an applied risk analysis of the results along several dimensions. Overall, the pilot demonstrated that the B3 dataset offers a viable, nuanced method for rapidly assessing the biosecurity risk posed by a LLM, identifying the key sources of that risk and providing guidance for priority areas of mitigation priority.

Multi-domain performance analysis with scores tailored to user preferences

arXiv:2512.08715v1 Announce Type: cross Abstract: The performance of algorithms, methods, and models tends to depend heavily on the distribution of cases on which they are applied, this distribution being specific to the applicative domain. After performing an evaluation in several domains, it is highly informative to compute a (weighted) mean performance and, as shown in this paper, to scrutinize what happens during this averaging. To achieve this goal, we adopt a probabilistic framework and consider a performance as a probability measure (e.g., a normalized confusion matrix for a classification task). It appears that the corresponding weighted mean is known to be the summarization, and that only some remarkable scores assign to the summarized performance a value equal to a weighted arithmetic mean of the values assigned to the domain-specific performances. These scores include the family of ranking scores, a continuum parameterized by user preferences, and that the weights to consider in the arithmetic mean depend on the user preferences. Based on this, we rigorously define four domains, named easiest, most difficult, preponderant, and bottleneck domains, as functions of user preferences. After establishing the theory in a general setting, regardless of the task, we develop new visual tools for two-class classification.

AI-powered virtual tissues from spatial proteomics for clinical diagnostics and biomedical discovery

arXiv:2501.06039v2 Announce Type: replace-cross Abstract: Spatial proteomics technologies have transformed our understanding of complex tissue architecture in cancer but present unique challenges for computational analysis. Each study uses a different marker panel and protocol, and most methods are tailored to single cohorts, which limits knowledge transfer and robust biomarker discovery. Here we present Virtual Tissues (VirTues), a general-purpose foundation model for spatial proteomics that learns marker-aware, multi-scale representations of proteins, cells, niches and tissues directly from multiplex imaging data. From a single pretrained backbone, VirTues supports marker reconstruction, cell typing and niche annotation, spatial biomarker discovery, and patient stratification, including zero-shot annotation across heterogeneous panels and datasets. In triple-negative breast cancer, VirTues-derived biomarkers predict anti-PD-L1 chemo-immunotherapy response and stratify disease-free survival in an independent cohort, outperforming state-of-the-art biomarkers derived from the same datasets and current clinical stratification schemes.

OMNIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and Modalities

arXiv:2505.23856v2 Announce Type: replace-cross Abstract: The emerging capabilities of large language models (LLMs) have sparked concerns about their immediate potential for harmful misuse. The core approach to mitigate these concerns is the detection of harmful queries to the model. Current detection approaches are fallible, and are particularly susceptible to attacks that exploit mismatched generalization of model capabilities (e.g., prompts in low-resource languages or prompts provided in non-text modalities such as image and audio). To tackle this challenge, we propose Omniguard, an approach for detecting harmful prompts across languages and modalities. Our approach (i) identifies internal representations of an LLM/MLLM that are aligned across languages or modalities and then (ii) uses them to build a language-agnostic or modality-agnostic classifier for detecting harmful prompts. Omniguard improves harmful prompt classification accuracy by 11.57\% over the strongest baseline in a multilingual setting, by 20.44\% for image-based prompts, and sets a new SOTA for audio-based prompts. By repurposing embeddings computed during generation, Omniguard is also very efficient ($\approx\!120 \times$ faster than the next fastest baseline). Code and data are available at: https://github.com/vsahil/OmniGuard.

Development of a Hospital-at-Home Digital Twin for Patients With Frailty: Scoping Review

Background: Increasing demand on healthcare systems requires innovative and transformative solutions to deliver efficient, high-quality care. One promising approach is Digital Twin (DT) technology, which leverages real time data to create dynamic virtual representations of a physical entity (individuals or space) to anticipate future scenarios and support care decisions. While DTs have been explored in various sectors, their application in Hospital at Home (HaH), which delivers acute level care in home environments, remains unexplored. Objective: This review bridges a critical knowledge gap and examines the existing evidence on DT-enabling tools for managing patients with frailty in home settings. This will identify the underpinning architectural components required to inform a HaH-DT system which can support clinical decision-making. Methods: Six electronic databases (Embase, MEDLINE, Cochrane CENTRAL, CINAHL, Web of Science and Scopus) were searched, along with grey literature, to identifying primary studies published in English, between January 2019 and September 2025. Included studies had to report on the monitoring or management of patients with frailty within their own home, and information was charted on a pre-defined data collection form to answer the research objectives. Review articles, protocols, and conference abstracts were excluded. Results: Sixty-nine reports were included, of which 54% (n=37) used quantitative approaches, and 36% (n=25) were pilot or feasibility studies. Reports were analysed for DT-enabling tools and systematically mapped across the proposed five-layered DT architecture: sensing, communication, storage, analytics, and visualisation. Taxonomies of DT layers, their interconnections, and the classifications of the types of data collected (e.g., about the patient, the home environment, the use of medical equipment) are presented. This evidence identifies DT-enabling tools used for a variety of functions and a range of sensing technologies that exist (e.g., passive sensing via wearables, active physiological sensors, ambient sensors to detect motion/environmental changes). The most prevalent modes of communication were wireless and network-based (n=36), with the majority using Bluetooth (n=12). This review highlights better understanding of data management, in particular secure storage, is required within local healthcare systems. The emerging potential of predictive and prescriptive analytics, which can enable clinicians to predict risk, support clinical decision-making, or activate alert-triggered health interventions were mapped. Existing evidence suggests analytics methods are currently largely descriptive with a lack of advanced methods such as prescriptive analytics to enable recommendations of an optimal course of action, and the absence of diagnostic analytics which can highlight why a situation has occurred. Reported DT-enabling tools demonstrate patient-centered benefits, including enhanced motivation, reassurance, and personalised care. However, concerns persist regarding device accuracy, user acceptability, and implications for carers and organisational workflows. Conclusions: This review is among the first to systematically map DT-enabling tools to inform a potential HaH-DT in patients with frailty and organised by a 5-layered conceptual model. Understanding these architectural layers provides the foundations to enable stakeholders advance research and development in areas where there are knowledge gaps and consider how a HaH DT can effectively operate within current healthcare systems. By leveraging technology-enabled care in complex home-based settings, there is great potential to deliver safer, personalised and timely care.

Somatic evolution following cancer treatment in normal tissue

Nature, Published online: 10 December 2025; doi:10.1038/s41586-025-09792-4

High-depth sequencing of non-cancerous tissue from patients with metastatic cancer reveals single-base mutational signatures of alcohol, smoking and cancer treatments, and reveals how exogenous factors, including cancer therapies, affect somatic cell evolution.
  • ✇AI News
  • Inside the playbook of companies winning with AI Muhammad Zulhusni
    Many companies are still working out how to use AI in a steady and practical way, but a small group is already pulling ahead. New research from NTT DATA outlines a playbook that shows how these “AI leaders” set themselves apart through strong plans, firm decisions, and a disciplined approach to building and using AI across their organisations. The findings come from a survey of 2,567 senior executives in 35 countries and 15 industries. Only 15% of the organisations met the bar to be considere
     

Inside the playbook of companies winning with AI

10 December 2025 at 17:00

Many companies are still working out how to use AI in a steady and practical way, but a small group is already pulling ahead. New research from NTT DATA outlines a playbook that shows how these “AI leaders” set themselves apart through strong plans, firm decisions, and a disciplined approach to building and using AI across their organisations.

The findings come from a survey of 2,567 senior executives in 35 countries and 15 industries. Only 15% of the organisations met the bar to be considered AI leaders. These companies share a few traits: clear direction on where AI fits into their business, a solid operating model, and consistent follow-through. They also reported higher revenue growth and stronger profit margins than everyone else in the study.

Yutaka Sasaki, President and CEO of NTT DATA Group, put it simply: “AI accountability now belongs in the boardroom and demands an enterprise-wide agenda. Our research shows that a small group of AI leaders already are using AI to differentiate, grow and reinvent how humans and machines create value together.”

The playbook behind strong AI plans

One of the clearest differences between leaders and the rest is how they approach strategy. For these companies, AI is not a side project or a tool bolted onto existing work. They treat it as a core driver of growth and adjust their plans to match that view.

A major advantage for these leaders is how closely they connect AI with their business goals. This alignment helps them move faster and stay focused, which in turn delivers stronger financial outcomes. They also zero in on a few high-value areas of the business rather than spreading resources too thin. By redesigning entire workflows around AI, they unlock more value than if they had only made small improvements in scattered parts of the organisation.

The report describes this as a kind of flywheel: early investments bring early wins, which then encourage more investment. Over time, this cycle becomes self-reinforcing. Leaders also rebuild important applications with AI embedded inside them, instead of adding basic AI features on top of old systems. This approach helps them see deeper impact and prepares the organisation for long-term gains.

How leaders put their plans to work

A good plan only works when backed by strong execution. AI leaders stand out through the foundations they build, the way they support their people, and how they drive adoption across the entire organisation.

These companies invest in secure and scalable systems that can support large AI workloads. In some cases, they shift or localise their infrastructure to support private or sovereign AI needs. They also work to remove system bottlenecks so teams can move without roadblocks.

Rather than using AI as a replacement for workers, leaders use it to help experienced employees do higher-value work. This “expert-first” approach allows teams to use their judgment while letting AI handle complex or time-consuming tasks.

AI leaders also focus on adoption as a long-term change effort. They treat it as a company-wide shift, supported by clear communication and structured change management. This helps reduce pushback and encourages steady use of AI at all levels.

Governance is another major difference. Leading organisations centralise their AI oversight, give clear responsibility to senior roles such as Chief AI Officers, and build processes that help balance innovation with risk. These systems allow them to scale AI more confidently.

Partnerships also play a major role. Top companies often bring in outside experts and are open to arrangements that tie outcomes to shared success. This helps them move faster while keeping their goals in view.

Abhijit Dubey, CEO and CAIO of NTT DATA, Inc., summarised the path forward: “Once AI and business strategies are aligned, the single most effective move is to pick one or two domains that deliver disproportionate value and redesign them end-to-end with AI. Supporting this focused, end-to-end approach with strong governance, modern infrastructure and trusted partners is how today’s AI leaders are turning pilots into profit and pulling ahead of the market.”

(Photo by Igor Omilaev)

See also: OpenAI: Enterprise users swap AI pilots for deep integrations

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Inside the playbook of companies winning with AI appeared first on AI News.

  • ✇STAT
  • STAT+: Pharmalittle: We’re reading about FDA plans for CAR-T therapies, skinny drug labels, and much more Ed Silverman
    Rise and shine, everyone. The middle of the week is upon us. Have heart, though. You made it this far, so why not hang on for another couple of days, yes? And what better way to make the time fly than to keep busy. So grab that cup of stimulation — our flavor today boasts the aroma of blueberries — and get started. Meanwhile, do keep us in mind if you hear anything interesting. Have a smashing day… In a closely watched case, the U.S. solicitor general urged the Supreme Court to review a contr
     

STAT+: Pharmalittle: We’re reading about FDA plans for CAR-T therapies, skinny drug labels, and much more

10 December 2025 at 22:27

Rise and shine, everyone. The middle of the week is upon us. Have heart, though. You made it this far, so why not hang on for another couple of days, yes? And what better way to make the time fly than to keep busy. So grab that cup of stimulation — our flavor today boasts the aroma of blueberries — and get started. Meanwhile, do keep us in mind if you hear anything interesting. Have a smashing day…

In a closely watched case, the U.S. solicitor general urged the Supreme Court to review a controversy over so-called skinny labels for medicines, arguing that an appeals court finding threatens the availability of lower-cost generic drugs, STAT tells us. Skinny labeling refers to a process in which a generic drug company seeks regulatory approval to market its medicine for a specific use, but not other patented uses for which a brand-name drug is prescribed. For instance, a generic drug could be marketed to treat one type of heart problem, but not another. In doing so, the generic company seeks to avoid lawsuits claiming patent infringement. Doubts were raised about the maneuver, however, when the Supreme Court two years ago declined to hear an appeal of a lower court ruling, which questioned the practice. Now, this second case is being seen as a test for whether skinny labeling can survive as a way for generic companies to market medicines.

The U.S. Food and Drug Administration is on track to make it harder for CAR-T therapy developers to bring their products to market by making full randomized, controlled trials the new standard it will accept for regulatory filings, Pharmaphorum writes. At the moment, it has been possible to develop CAR-Ts based on single-arm trials, although some have used an active comparator. Now, with the number of CAR-Ts on the market now in double figures, the FDA is eyeing RCTs with a control group as well as “a survival or acceptable time-to-event endpoint.” The move towards a higher threshold for showing efficacy for new CAR-Ts comes after the FDA loosened requirements for safety monitoring by eliminating the risk evaluation and mitigation strategies previously required for already-marketed therapies targeting CD19 and BCMA, which the agency said would make them more accessible.

Continue to STAT+ to read the full story…

© Alex Hogan/STAT

KidSpeak: A General Multi-purpose LLM for Kids' Speech Recognition and Screening

arXiv:2512.05994v1 Announce Type: cross Abstract: With the rapid advancement of conversational and diffusion-based AI, there is a growing adoption of AI in educational services, ranging from grading and assessment tools to personalized learning systems that provide targeted support for students. However, this adaptability has yet to fully extend to the domain of children's speech, where existing models often fail due to their reliance on datasets designed for clear, articulate adult speech. Children, particularly those in early developmental stages or with speech and language pathologies, present unique challenges that current AI models and datasets are ill-equipped to handle. To address this, we introduce KidSpeak, a multi-task speech-enhanced Foundation Model capable of both generative and discriminative tasks specifically tailored to children's speech patterns. Our framework employs a two-stage training process that incorporates phonetic knowledge into the speech encoder, achieving an average accuracy of 87% across four separate tasks. Furthermore, recognizing the limitations of scalable human annotation and existing speech alignment tools, we propose the Flexible and Automatic Speech Aligner (FASA) and leverage the method to construct high quality datasets for training and evaluation. This novel alignment tool significantly improves the quality of aligned children's speech from noisy data, enhancing data quality by 13.6x compared to human annotations, as demonstrated on the CHILDES dataset. To the best of our knowledge, KidSpeak and FASA represent the first comprehensive solution designed for speech and language therapy in children, offering both a multi-purpose speech LLM and a robust alignment tool.

Data Visualization Support for Interdisciplinary Team Treatment Planning in Clinical Oncology: Scoping Review

Background: Complex and expanding datasets in clinical oncology applications require flexible and interactive visualization of patient data to provide physicians and other medical professionals with maximum amount of information. In particular, interdisciplinary tumor conferences profit from customized tools to integrate, link, and visualize relevant data from all professions involved. Objective: Our objective was to identify and present currently available data visualization tools for tumor boards and related areas. We wanted to provide an overview of not only the digital tools currently used in tumor board settings but also of the data they include, their respective visualization solutions, and their integration into hospital processes. Methods: This scoping review was based on the scoping study framework by Arksey and O’Malley and attempted to answer the following research question: “What are the key features of data visualization solutions used in molecular and organ tumor boards, and how are these elements integrated and used within the clinical setting?” The following electronic databases were searched for articles: PubMed, Web of Science, and Scopus. Articles were deemed eligible if published in English in the last 10 years. Eligible articles were first deduplicated, followed by screening of titles and abstracts. Full-text screening was then conducted to decide on article selection. All included articles were analyzed using a data extraction template. The template included a variety of meta-information, as well as specific fields aiming to answer the research question. Results: The review process started with 2049 articles, of which 1014 (49.49%) were included in the title and abstract screening. A total of 5.47% (112/2049) of the publications were eligible for full-text screening, leading to 2.93% (60/2049) of the publications being eligible for final inclusion. They covered 49 distinct visualization tools and applications. We discovered a variety of innovative visualization solutions, most often driven by the complexity of omics data, represented in 96% (47/49) of the tools. Tables remained the most used tool for the visualization of data types described in the articles. Approximately one-third of the identified tools (16/49, 33%) were systematically evaluated in some form. For most discovered tools (37/49, 76%), there was no documentation of implementation into the clinical routine. A significant number of applications (21/49, 43%) were available through open-source access. Conclusions: There is a wide range of projects providing visualization solutions for tumor boards and clinical oncology applications. Among the few tools that have made their way into clinical routine settings, there are both commercial and academic solutions. While tables for a variety of data types remain the dominant visualization strategy, the complexity of omics data appears to be the driving force behind many visualization innovations in the domain of tumor boards. Trial Registration:

WisPaper: Your AI Scholar Search Engine

arXiv:2512.06879v1 Announce Type: cross Abstract: Researchers struggle to efficiently locate and manage relevant literature within the exponentially growing body of scientific publications. We present \textsc{WisPaper}, an intelligent academic retrieval and literature management platform that addresses this challenge through three integrated capabilities: (1) \textit{Scholar Search}, featuring both quick keyword-based and deep agentic search modes for efficient paper discovery; (2) \textit{Library}, a customizable knowledge base for systematic literature organization; and (3) \textit{AI Feeds}, an intelligent recommendation system that automatically delivers relevant new publications based on user interests. Unlike existing academic tools, \textsc{WisPaper} provides a closed-loop workflow that seamlessly connects literature discovery, management, and continuous tracking of research frontiers. Our multilingual and multidisciplinary system significantly reduces the time researchers from diverse backgrounds spend on paper screening and management, enabling them to focus on their core research activities. The platform is publicly accessible and serves researchers across academia and industry.

A Field Guide to Deploying AI Agents in Clinical Practice

arXiv:2509.26153v3 Announce Type: replace Abstract: Large language models (LLMs) integrated into agent-driven workflows hold immense promise for healthcare, yet a significant gap exists between their potential and practical implementation within clinical settings. To address this, we present a practitioner-oriented field manual for deploying generative agents that use electronic health record (EHR) data. This guide is informed by our experience deploying the "irAE-Agent", an automated system to detect immune-related adverse events from clinical notes at Mass General Brigham, and by structured interviews with 21 clinicians, engineers, and informatics leaders involved in the project. Our analysis reveals a critical misalignment in clinical AI development: less than 20% of our effort was dedicated to prompt engineering and model development, while over 80% was consumed by the sociotechnical work of implementation. We distill this effort into five "heavy lifts": data integration, model validation, ensuring economic value, managing system drift, and governance. By providing actionable solutions for each of these challenges, this field manual shifts the focus from algorithmic development to the essential infrastructure and implementation work required to bridge the "valley of death" and successfully translate generative AI from pilot projects into routine clinical care.

Algorithms Trained on Normal Chest X-rays Can Predict Health Insurance Types

arXiv:2511.11030v4 Announce Type: replace-cross Abstract: Artificial intelligence is revealing what medicine never intended to encode. Deep vision models, trained on chest X-rays, can now detect not only disease but also invisible traces of social inequality. In this study, we show that state-of-the-art architectures (DenseNet121, SwinV2-B, MedMamba) can predict a patient's health insurance type, a strong proxy for socioeconomic status, from normal chest X-rays with significant accuracy (AUC around 0.70 on MIMIC-CXR-JPG, 0.68 on CheXpert). The signal was unlikely contributed by demographic features by our machine learning study combining age, race, and sex labels to predict health insurance types; it also remains detectable when the model is trained exclusively on a single racial group. Patch-based occlusion reveals that the signal is diffuse rather than localized, embedded in the upper and mid-thoracic regions. This suggests that deep networks may be internalizing subtle traces of clinical environments, equipment differences, or care pathways; learning socioeconomic segregation itself. These findings challenge the assumption that medical images are neutral biological data. By uncovering how models perceive and exploit these hidden social signatures, this work reframes fairness in medical AI: the goal is no longer only to balance datasets or adjust thresholds, but to interrogate and disentangle the social fingerprints embedded in clinical data itself.
❌