❌

Normal view

Step into the future: The full AI Stage agenda at TechCrunch Disrupt 2025

24 September 2025 at 22:30
The AI Stage at TechCrunch Disrupt 2025 is officially locked and loaded, featuring the powerhouses shaping the future of artificial intelligence.

Expanding care coordination in an integrated health system through causal machine learning

npj Digital Medicine, Published online: 24 September 2025; doi:10.1038/s41746-025-01925-3

Expanding care coordination in an integrated health system through causal machine learning

Article: InfoQ AI, ML and Data Engineering Trends Report - 2025

This InfoQ Trends Report offers readers a comprehensive overview of emerging trends and technologies in the areas of AI, ML, and Data Engineering. This report summarizes the InfoQ editorial team’s and external guests' view on the current trends in AI and ML technologies and what to look out for in the next 12 months.

By Srini Penchikala, Savannah Kunovsky, Anthony Alford, Daniel Dominguez, Vinod Goje

Fine-Tuning Methods for Large Language Models in Clinical Medicine by Supervised Fine-Tuning and Direct Preference Optimization: Comparative Evaluation

Background: Large language model (LLM) fine tuning is the process of adjusting out-of-the-box model weights using a dataset of interest. Fine tuning can be a powerful technique to improve model performance in fields like medicine, where data access is restricted and LLMs may have poor out-of-the-box performance. Objective: In this study we investigated the benefits of fine tuning with supervised fine tuning (SFT) and direct preference optimization (DPO) across a range of LLM applications for medicine Methods: We use Llama3 7B and Mistral 7B v2 to compare the performance of SFT and DPO across four datasets for common natural language tasks in medicine. The tasks evaluated were simple classification, clinical reasoning, summarization, and clinical triage. Results: Clinical Reasoning accuracy increased 8% and 7% with DPO over SFT for Llama3 (p value 0.003) and Mistral2 (p value 0.004) respectively. Summarization quality, graded on a five point Likert scale, increased 0.13 and 0.10 for Llama3 and Mistral2 (p values

Comparative Evaluation of a Medical Large Language Model in Answering Real-World Radiation Oncology Questions: Multicenter Observational Study

Background: Large language models (LLMs) hold promise for supporting clinical tasks, particularly in data-driven and technical disciplines such as radiation oncology. While prior evaluation studies have focused on examination-style settings for evaluating LLMs, their performance in real-life clinical scenarios remains unclear. In the future, LLMs might be used as general AI assistants to answer questions arising in clinical practice. It is unclear how well a modern LLM, locally executed within the infrastructure of a hospital, would answer such questions compared with clinical experts. Objective: This study aimed to assess the performance of a locally deployed, state-of-the-art medical LLM in answering real-world clinical questions in radiation oncology compared with clinical experts. The aim was to evaluate the overall quality of answers, as well as the potential harmfulness of the answers if used for clinical decision-making. Methods: Physicians from 10 departments of European hospitals collected questions arising in the clinical practice of radiation oncology. Fifty of these questions were answered by 3 senior radiation oncology experts with at least 10 years of work experience, as well as the LLM Llama3-OpenBioLLM-70B (Ankit Pal and Malaikannan Sankarasubbu). In a blinded review, physicians rated the overall answer quality on a 5-point Likert scale (quality), assessed whether an answer might be potentially harmful if used for clinical decision-making (harmfulness), and determined if responses were from an expert or the LLM (recognizability). Comparisons between clinical experts and LLMs were then made for quality, harmfulness, and recognizability. Results: There were no significant differences between the quality of the answers between LLM and clinical experts (mean scores of 3.38 vs 3.63; median 4.00, IQR 3.00-4.00 vs median 3.67, IQR 3.33-4.00; P=.26; Wilcoxon signed rank test). The answers were deemed potentially harmful in 13% of cases for the clinical experts compared with 16% of cases for the LLM (P=.63; Fisher exact test). Physicians correctly identified whether an answer was given by a clinical expert or an LLM in 78% and 72% of cases, respectively. Conclusions: A state-of-the-art medical LLM can answer real-life questions from the clinical practice of radiation oncology similarly well as clinical experts regarding overall quality and potential harmfulness. Such LLMs can already be deployed within the local hospital environment at an affordable cost. While LLMs may not yet be ready for clinical implementation as general AI assistants, the technology continues to improve at a rapid pace. Evaluation studies based on real-life situations are important to better understand the weaknesses and limitations of LLMs in clinical practice. Such studies are also crucial to define when the technology is ready for clinical implementation. Furthermore, education for health care professionals on generative AI is needed to ensure responsible clinical implementation of this transforming technology.

Circulating tumor DNA in patients with cancer: insights from clinical laboratory

Adv Lab Med. 2025 Jun 16;6(3):259-276. doi: 10.1515/almed-2025-0010. eCollection 2025 Sep.

ABSTRACT

Blood-based circulating tumor DNA (ctDNA) analysis has emerged as a highly relevant non-invasive method for molecular profiling of solid tumors, offering valuable information about the genetic landscape of cancer. Somatic mutation analysis of ctDNA is now used clinically to guide targeted therapies for advanced cancers. Recent advancements have also revealed its potential in early detection, prognosis, minimal residual disease assessment, and prediction/monitoring of therapeutic response. In recent years, significant progress has been made with the development of various PCR and NGS-based methods designed for assessing gene variants in ctDNA of patients with cancer. However, despite the transformative possibilities that ctDNA analysis presents, challenges persist. Standardization of preanalytical and analytical protocols, assay sensitivity, and the interpretation of results remain critical hurdles that need to be addressed for the widespread clinical implementation of ctDNA testing. In addition to somatic mutations, emerging studies on DNA methylation (epigenomics) and fragment size patterns (fragmentomics) in several types of biological fluids are yielding promising results as non-invasive biomarkers for effective cancer management. This review addresses the clinical applications of somatic gene variants in ctDNA, emphasizes their potential as cancer biomarkers, and highlights essential factors for successful implementation in clinical laboratories and cancer management.

PMID:40977813 | PMC:PMC12446922 | DOI:10.1515/almed-2025-0010

A statistical physics approach to integrating multi-omics data for disease-module detection

Cell Rep Methods. 2025 Sep 19:101183. doi: 10.1016/j.crmeth.2025.101183. Online ahead of print.

ABSTRACT

Genes associated with the same disease frequently engage in mutual biological interactions, e.g., perturbation within a specific neighborhood in the molecular interactome, often referred to as the disease module. This has propelled the advancement of network-based approaches toward elucidating the molecular bases of human diseases. Although many computational methods have been developed to integrate the molecular interactome and omics profiles to extract such context-dependent disease modules, approaches that leverage multi-omics for disease-module detection are still lacking. Here, we developed a statistical physics approach based on the random-field O(n) model (RFOnM) to fill this gap. We applied the RFOnM approach to integrate gene-expression data and genome-wide association studies or mRNA data and DNA methylation for several complex diseases with the human interactome. We found that the RFOnM approach outperforms existing single omics methods in most of the complex diseases considered in this study.

PMID:40975055 | DOI:10.1016/j.crmeth.2025.101183

Implementation of a Virtual Hospital in the Home Service for Patients With COVID-19 in Queensland, Australia: Mixed Methods Evaluation Using the RE-AIM Framework

Background: Hospital in the home (HITH) provides home-based care as an alternative to traditional hospitalization. In response to the COVID-19 Omicron wave, a public hospital in the rural Western portion of Southeast Queensland implemented a virtual HITH service to support adults, maternity patients, and children with moderate COVID-19 symptoms and additional health concerns. Although the pandemic accelerated the uptake of virtual care within HITH models, existing literature has focused on clinical outcomes, with limited evidence on key implementation outcomes. Objective: Using the RE-AIM (reach, effectiveness, adoption, implementation, and maintenance) framework, this study evaluated the implementation of the virtual COVID-19 HITH service and identified factors influencing its implementation, to inform ongoing service development and support potential scaling of this model of care. Methods: The RE-AIM implementation science framework was selected to guide the evaluation, capturing both clinical and contextual dimensions of implementation at both individual and organizational levels. Quantitative data on service usage and costs were retrospectively extracted from electronic medical records and finance records, while patient experience data were drawn from patient-reported experience measures surveys. Qualitative data were collected through one-on-one interviews with patients and staff. All data sources were analyzed separately and then triangulated within the RE-AIM framework to understand what occurred, how, and why. Results: The service admitted 3192 patients, most of whom were female (2027/3192, 63.5%), English-speaking (3140/3192, 98.4%), and residing in socioeconomically disadvantaged areas (1879/3192, 58.9%) (reach). The model was feasible and safe to implement, managing 3240 admissions with no reported deaths. Patients valued continuous access to care and described better recovery experiences at home (effectiveness). Staff viewed the model as appropriate for identifying and managing high-risk patients in the community, easing pressure on hospital beds (adoption). The service cost Aus $ 5.4 million (US $3.5 million) over 11 months. Implementation barriers included the urgency of the pandemic scenario, limited infrastructure and human resources, and changing requirements in relation to COVID-19. These were mitigated by several people factors that were critical to its successful implementation, including a consultant-led structure, staff commitment, and adaptability (implementation). The service saved 16,651 inpatient bed days before being integrated into core HITH operations. The experience strengthened staff capabilities in emergency response, virtual care delivery, and strategic planning. The model shows promise for broader application into pediatric care, though further work is needed to enhance interdepartmental collaboration and staff recognition (maintenance). Conclusions: This study demonstrated that a virtual HITH model can be implemented effectively and safely at scale. Findings support its potential for integration into routine care, provided that adequate resource planning, a skilled and multidisciplinary workforce, well-defined care pathways, and equity-focused strategies are in place.

Opinion: Four reasons why generative AI chatbots could lead to psychosis in vulnerable people

18 September 2025 at 16:30

Three scholars discovered a strange mirror deep in the forest. It spoke to them in a soothing voice and answered all their questions warmly, knowledgeably, and eloquently.

The captivated scholars became obsessed, whispering one secret after another to the mirror. It replied with affection, promise, and meaning that kept them returning to it. They began ignoring one another, each convinced the mirror “understood” them best.

Read the rest…

© Adobe

From frameworks to finance: how sharing benefits from the use of digital sequence information can evolve to contribute to biodiversity conservation

Nature Biotechnology, Published online: 18 September 2025; doi:10.1038/s41587-025-02820-8

The COP16 decision established a multilateral mechanism for digital sequence information (DSI) benefit-sharing. This Comment brings together insights from academia and commercial DSI researchers to assess what has been accomplished so far, identify remaining challenges and describe elements under discussion to support collective goals.

Diagnostic Performance of Computed Tomography–Based Artificial Intelligence for Early Recurrence of Cholangiocarcinoma: Systematic Review and Meta-Analysis

Background: Despite artificial intelligence (AI) models demonstrating high predictive accuracy for early cholangiocarcinoma recurrence, their clinical application faces challenges, such as reproducibility, generalizability, hidden biases, and uncertain performance across diverse datasets and populations, raising concerns about their practical applicability. Objective: This meta-analysis aims to systematically assess the diagnostic performance of AI models using computed tomography (CT) imaging to predict early recurrence of cholangiocarcinoma. Methods: A systematic search was conducted in PubMed, Embase, and Web of Science for studies published up to May 2025. Studies were selected based on the Participants, Index test, Target condition, Reference standard, Outcomes, and Setting (PITROS) framework. Participants included patients diagnosed with cholangiocarcinoma (including intrahepatic and extrahepatic locations). The index test was AI techniques applied to CT imaging for early recurrence prediction (defined as within 1 year), while the target condition was early recurrence of cholangiocarcinoma (positive group: recurrence; negative group: no recurrence). The reference standard was pathological diagnosis or imaging follow-up confirming recurrence. Outcomes included sensitivity, specificity, diagnostic odds ratio (DOR), and area under the receiver operating characteristic curve (AUC), assessed in both internal and external validation cohorts. The setting comprised retrospective or prospective studies using hospital datasets. Methodological quality was assessed using an optimized version of the revised Quality Assessment of Diagnostic Accuracy Studies-2 tool. Heterogeneity was assessed using the I² statistic. Pooled sensitivity, specificity, DOR, and AUC were calculated using a bivariate random-effects model. Results: A total of 9 studies with 30 datasets involving 1537 patients were included. In internal validation cohorts, CT-based AI models showed a pooled sensitivity of 0.87 (95% CI 0.81-0.92), specificity of 0.85 (95% CI 0.79-0.89), DOR of 37.71 (95% CI 18.35-77.51), and AUC of 0.93 (95% CI 0.90-0.94). In external validation cohorts, pooled sensitivity was 0.87 (95% CI 0.81-0.91), specificity was 0.82 (95% CI 0.77-0.86), DOR was 30.81 (95% CI 18.79-50.52), and AUC was 0.85 (95% CI 0.82-0.88). The AUC was significantly lower in external validation cohorts compared to internal validation cohorts (P<.001). Conclusions: Our results show that CT-based AI models predict early cholangiocarcinoma recurrence with high performance in internal validation sets and moderate performance in external validation sets. However, the high heterogeneity observed may impact the robustness of these results. Future research should focus on prospective studies and establishing standardized gold standards to further validate the clinical applicability and generalizability of AI models.

Large Language Models’ Clinical Decision-Making on When to Perform a Kidney Biopsy: Comparative Study

Background: Artificial intelligence (AI) and Large Language models (LLMs) are increasing in sophistication and are being integrated into many disciplines. The potential for LLMs to augment clinical decisions is an evolving area of research. Objective: This study compared the responses of over 1000 kidney specialist physicians (nephrologists) to outputs of commonly used LLMs using a questionnaire determining when a kidney biopsy should be performed. Methods: This research group completed a large online questionnaire for nephrologists to determine when a kidney biopsy should be performed. The questionnaire was co-designed with patient participation, refined through multiple iterations, then piloted locally before international dissemination. It was the largest international study in the field and demonstrated variation between human clinicians in biopsy propensity relating to human factors such as sex and age, as well as systemic factors such as country, job seniority and technical proficiency. The same questions were put to both human doctors and LLMs in an identical order in a single session. Eight commonly used LLMs were interrogated: Chat GPT 3.5, Mistral Hugging Face, Perplexity, Microsoft Co-pilot, Llama 2, GPT 4.0, MedLM and Claude 3. The most common response given by clinicians (human mode) to each question was taken as the baseline for comparison. Questionnaire responses to the indications and contraindications for biopsy generated a score (0-44) reflecting biopsy propensity, in which a higher score was used as a surrogate marker for an increased tolerance of potential associated risks. Results: The ability of LLMs to reproduce human expert consensus varied widely with some models demonstrating a balanced approach to risk in a similar manner to humans, whilst other models reported outputs at either end of the spectrum for risk tolerance. In terms of agreement with the human mode, Chat GPT 3.5 and GPT 4.0 (Open AI) had the highest levels of alignment, with the human mode selected in 6/11 questions. The total biopsy propensity score generated from the human mode was 23/44. Both Open AI models produced similar propensity scores between 22 and 24, however Llama 2 and MS Co-pilot also reported scores within this range, but with poorer response alignment to the human mode at only 2/11 questions. The most risk averse model in this study was MedLM with a propensity score of 11 and the least risk averse model was Claude 3 with a score of 34. Conclusions: LLM outputs demonstrated a modest ability to replicate human clinical decision making in this study, however the performance varied widely between LLM models. Questions with more uniform human responses produced LLM outputs with greater alignment, whereas in questions with low levels of human consensus there was poor output alignment. This may limit the practical use of LLMs in real world clinical practice.

Navigating the Boundaries of Teleconsultation—Capabilities, Limitations, and Pathways for Improvement: Qualitative Study of the Experiences of Patients With Stroke

Background: Survivors of stroke often face persistent challenges accessing postdischarge care due to mobility limitations, transportation burdens, and inflexible scheduling. Teleconsultation has emerged as a potential solution to improve continuity of care, but its perceived strengths and limitations from the patient perspective remain insufficiently understood. Objective: This study aimed to explore the experiences of survivors of stroke with a nurse-led teleconsultation program to (1) identify perceived capabilities; (2) understand limitations in usability, accessibility, and clinical function; and (3) generate patient-informed recommendations for improvement. Methods: A qualitative study was embedded within a 3-month nurse-led teleconsultation intervention delivered by advanced practice nurses. A total of 21 survivors of ischemic stroke (aged 45-76 y; female: n=11, 52%) who had preserved cognitive function (Montreal Cognitive Assessment score ≥22) and smartphone access participated in 6 focus groups conducted via Zoom. Data were analyzed thematically using an established framework. Data saturation was achieved. Results: Participants widely valued teleconsultation for reducing logistical burdens; enhancing access; and offering a more comfortable, emotionally supportive setting for follow-up care. Many reported increased awareness and motivation for self-monitoring. However, limitations included an inability to perform physical assessments or respond to emergencies; digital and usability barriers, especially among older users; and scheduling inflexibility. Participants emphasized the need for patient-initiated follow-up mechanisms, physician collaboration for medication management, and greater support for users considered digitally marginalized. They also highlighted the potential of teleconsultation to serve as a triage tool, reserving in-person care for complex cases. Conclusions: Nurse-led teleconsultation was perceived as a convenient and supportive modality for poststroke care, particularly for stable follow-ups and psychosocial support. However, its long-term viability depends on addressing clinical and technical limitations, enhancing user autonomy, and integrating interdisciplinary input. By centering the lived experiences of survivors of stroke, this study offers concrete recommendations to guide the development of more inclusive, responsive, and patient-centered teleconsultation models.

The arts for disease prevention and health promotion: a systematic review

Nature Medicine, Published online: 18 September 2025; doi:10.1038/s41591-025-03962-7

The arts, according to a systematic synthesis of data from 95 studies (across 26 countries), may support non-communicable disease prevention by providing opportunities for increased physical activity, and helping to address social forces that contribute to health inequities.

Bridging Technology and Pretest Genetic Services: Quantitative Study of Chatbot Interaction Patterns, User Characteristics, and Genetic Testing Decisions

Background: Among the alternative solutions being tested to improve access to genetic services, chatbots (or conversational agents) are being increasingly used for service delivery. Despite the growing number of studies on the accessibility and feasibility of chatbot genetic service delivery, limited attention has been paid to user interactions with chatbots in a real-world health care context. Objective: We examined users’ interaction patterns with a pretest cancer genetics education chatbot as well as the associations between users’ clinical and sociodemographic characteristics, chatbot interaction patterns, and genetic testing decisions. Methods: We analyzed data from the experimental arm of Broadening the Reach, Impact, and Delivery of Genetic Services, a multisite genetic services pragmatic trial in which participants eligible for hereditary cancer genetic testing based on family history were randomized to receive a chatbot intervention or standard care. In the experimental chatbot arm, participants were offered access to core educational content delivered by the chatbot with the option to select up to 9 supplementary informational prompts and ask open-ended questions. We computed descriptive statistics for the following interaction patterns: prompt selections, open-ended questions, completion status, dropout points, and postchat decisions regarding genetic testing. Logistic regression models were used to examine the relationships between clinical and sociodemographic factors and chatbot interaction variables, examining how these factors affected genetic testing decisions. Results: Of the 468 participants who initiated a chat, 391 (83.5%) completed it, with 315 (80.6%) of the completers expressing a willingness to pursue genetic testing. Of the 391 completers, 336 (85.9%) selected at least one informational prompt, 41 (10.5%) asked open-ended questions, and 3 (0.8%) opted for extra examples of risk information. Of the 77 noncompleters, 57 (74%) dropped out before accessing any informational content. Interaction patterns were not associated with clinical and sociodemographic factors except for prompt selection (varied by study site) and completion status (varied by family cancer history type). Participants who selected ≥3 prompts (odds ratio 0.33, 95% CI 0.12-0.91; P=.03) or asked open-ended questions (odds ratio 0.46, 95% CI 0.22-0.96; P=.04) were less likely to opt for genetic testing. Conclusions: Findings highlight the chatbot’s effectiveness in engaging users and its high acceptability, with most participants completing the chat, opting for additional information, and showing a high willingness to pursue genetic testing. Sociodemographic factors were not associated with interaction patterns, potentially indicating the chatbot’s scalability across diverse populations provided they have internet access. Future efforts should address the concerns of users with high information needs and integrate them into chatbot design to better support informed genetic decision-making.

Delegation to artificial intelligence can increase dishonest behaviour

Nature, Published online: 17 September 2025; doi:10.1038/s41586-025-09505-x

People cheat more when they delegate tasks to artificial intelligence, and large language models are more likely than humans to comply with unethical instructions—a risk that can be minimized by introducing prohibitive, task-specific guardrails.

New Doc on the Block: Scoping Review of AI Systems Delivering Motivational Interviewing for Health Behavior Change

Background: Artificial intelligence (AI) is increasingly used in digital health, particularly through large language models (LLMs), to support patient engagement and behavior change. One novel application is the delivery of motivational interviewing (MI), an evidence-based, patient-centered counseling technique designed to enhance motivation and resolve ambivalence around health behaviors. AI tools, including chatbots, mobile apps, and web-based agents, are being developed to simulate MI techniques at scale. While these innovations are promising, important questions remain about how faithfully AI systems can replicate MI principles or achieve meaningful behavioral impact. Objective: This scoping review aimed to summarize existing empirical studies evaluating AI-driven systems that apply MI techniques to support health behavior change. Specifically, we examined the feasibility of these systems; their fidelity to MI principles; and their reported behavioral, psychological, or engagement outcomes. Methods: We systematically searched PubMed, Embase, Scopus, Web of Science, and Cochrane Library for empirical studies published between January 1, 2018, and February 25, 2025. Eligible studies involved AI-driven systems using natural language generation, understanding, or computational logic to deliver MI techniques to users targeting a specific health behavior. We excluded studies using AI solely for training clinicians in MI. Three independent reviewers screened and extracted data on study design, AI modality and type, MI components, health behavior focus, MI fidelity assessment, and outcome domains. Results: Of the 1001 records identified, 15 (1.5%) met the inclusion criteria. Of these 15 studies, 6 (40%) were exploratory feasibility or pilot studies, and 3 (20%) were randomized controlled trials. AI modalities included rule-based chatbots (9/15, 60%), LLM-based systems (4/15, 27%), and virtual or mobile agents (2/15, 13%). Targeted behaviors included smoking cessation (6/15, 40%), substance use (3/15, 20%), COVID-19 vaccine hesitancy, type 2 diabetes self-management, stress, mental health service use, and opioid use during pregnancy. Of the 15 studies, 13 (87%) reported positive findings on feasibility or user acceptability, while 6 (40%) assessed MI fidelity using expert review or structured coding, with moderate to high alignment reported. Several studies found that users perceived the AI systems as judgment free, supportive, and easier to engage with than human counselors, particularly in stigmatized contexts. However, limitations in empathy, safety transparency, and emotional nuance were commonly noted. Only 3 (20%) of the 15 studies reported substantially significant behavioral changes. Conclusions: AI systems delivering MI show promise for enhancing patient engagement and scaling behavior change interventions. Early evidence supports their usability and partial fidelity to MI principles, especially in sensitive domains. However, most systems remain in early development, and few have been rigorously tested. Future research should prioritize randomized evaluations; standardized fidelity measures; and safeguards for LLM safety, empathy, and accuracy in health-related dialogue. Trial Registration: OSF Registries 10.17605/OSF.IO/G9N7E; https://osf.io/g9n7e
❌