❌

Normal view

  • ✇STAT
  • STAT+: Fresh data on hospital AI use & Califf dishes on tech Mario Aguilar
    You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday. Califf warns AI in health care ‘overhyped’ On a makeshift stage in a Midtown Manhattan office earlier this week,former Food and Drug Administration Commissioner Robert Califf struck a measured tone about the potential for artificial intelligence in health care. Asked whether the technology was o
     

STAT+: Fresh data on hospital AI use & Califf dishes on tech

18 September 2025 at 22:18

You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday.

Califf warns AI in health care ‘overhyped’

On a makeshift stage in a Midtown Manhattan office earlier this week,former Food and Drug Administration Commissioner Robert Califf struck a measured tone about the potential for artificial intelligence in health care. Asked whether the technology was overhyped he said it was. “I hear way too much about the money. I’m not hearing a lot of human values coming through discussions,” he said. Adding:

“Almost all of the technology is being applied to optimizing the financial status of healthcare delivery entities or companies that are making medical products and that’s not aligned with equitable, better patient outcomes. So until someone puts a soul back in the system, I think it’s going to get worse and worse.”

Continue to STAT+ to read the full story…

© Adobe

Opinion: Four reasons why generative AI chatbots could lead to psychosis in vulnerable people

18 September 2025 at 16:30

Three scholars discovered a strange mirror deep in the forest. It spoke to them in a soothing voice and answered all their questions warmly, knowledgeably, and eloquently.

The captivated scholars became obsessed, whispering one secret after another to the mirror. It replied with affection, promise, and meaning that kept them returning to it. They began ignoring one another, each convinced the mirror “understood” them best.

Read the rest…

© Adobe

From frameworks to finance: how sharing benefits from the use of digital sequence information can evolve to contribute to biodiversity conservation

Nature Biotechnology, Published online: 18 September 2025; doi:10.1038/s41587-025-02820-8

The COP16 decision established a multilateral mechanism for digital sequence information (DSI) benefit-sharing. This Comment brings together insights from academia and commercial DSI researchers to assess what has been accomplished so far, identify remaining challenges and describe elements under discussion to support collective goals.

Diagnostic Performance of Computed Tomography–Based Artificial Intelligence for Early Recurrence of Cholangiocarcinoma: Systematic Review and Meta-Analysis

Background: Despite artificial intelligence (AI) models demonstrating high predictive accuracy for early cholangiocarcinoma recurrence, their clinical application faces challenges, such as reproducibility, generalizability, hidden biases, and uncertain performance across diverse datasets and populations, raising concerns about their practical applicability. Objective: This meta-analysis aims to systematically assess the diagnostic performance of AI models using computed tomography (CT) imaging to predict early recurrence of cholangiocarcinoma. Methods: A systematic search was conducted in PubMed, Embase, and Web of Science for studies published up to May 2025. Studies were selected based on the Participants, Index test, Target condition, Reference standard, Outcomes, and Setting (PITROS) framework. Participants included patients diagnosed with cholangiocarcinoma (including intrahepatic and extrahepatic locations). The index test was AI techniques applied to CT imaging for early recurrence prediction (defined as within 1 year), while the target condition was early recurrence of cholangiocarcinoma (positive group: recurrence; negative group: no recurrence). The reference standard was pathological diagnosis or imaging follow-up confirming recurrence. Outcomes included sensitivity, specificity, diagnostic odds ratio (DOR), and area under the receiver operating characteristic curve (AUC), assessed in both internal and external validation cohorts. The setting comprised retrospective or prospective studies using hospital datasets. Methodological quality was assessed using an optimized version of the revised Quality Assessment of Diagnostic Accuracy Studies-2 tool. Heterogeneity was assessed using the I² statistic. Pooled sensitivity, specificity, DOR, and AUC were calculated using a bivariate random-effects model. Results: A total of 9 studies with 30 datasets involving 1537 patients were included. In internal validation cohorts, CT-based AI models showed a pooled sensitivity of 0.87 (95% CI 0.81-0.92), specificity of 0.85 (95% CI 0.79-0.89), DOR of 37.71 (95% CI 18.35-77.51), and AUC of 0.93 (95% CI 0.90-0.94). In external validation cohorts, pooled sensitivity was 0.87 (95% CI 0.81-0.91), specificity was 0.82 (95% CI 0.77-0.86), DOR was 30.81 (95% CI 18.79-50.52), and AUC was 0.85 (95% CI 0.82-0.88). The AUC was significantly lower in external validation cohorts compared to internal validation cohorts (P<.001). Conclusions: Our results show that CT-based AI models predict early cholangiocarcinoma recurrence with high performance in internal validation sets and moderate performance in external validation sets. However, the high heterogeneity observed may impact the robustness of these results. Future research should focus on prospective studies and establishing standardized gold standards to further validate the clinical applicability and generalizability of AI models.

Large Language Models’ Clinical Decision-Making on When to Perform a Kidney Biopsy: Comparative Study

Background: Artificial intelligence (AI) and Large Language models (LLMs) are increasing in sophistication and are being integrated into many disciplines. The potential for LLMs to augment clinical decisions is an evolving area of research. Objective: This study compared the responses of over 1000 kidney specialist physicians (nephrologists) to outputs of commonly used LLMs using a questionnaire determining when a kidney biopsy should be performed. Methods: This research group completed a large online questionnaire for nephrologists to determine when a kidney biopsy should be performed. The questionnaire was co-designed with patient participation, refined through multiple iterations, then piloted locally before international dissemination. It was the largest international study in the field and demonstrated variation between human clinicians in biopsy propensity relating to human factors such as sex and age, as well as systemic factors such as country, job seniority and technical proficiency. The same questions were put to both human doctors and LLMs in an identical order in a single session. Eight commonly used LLMs were interrogated: Chat GPT 3.5, Mistral Hugging Face, Perplexity, Microsoft Co-pilot, Llama 2, GPT 4.0, MedLM and Claude 3. The most common response given by clinicians (human mode) to each question was taken as the baseline for comparison. Questionnaire responses to the indications and contraindications for biopsy generated a score (0-44) reflecting biopsy propensity, in which a higher score was used as a surrogate marker for an increased tolerance of potential associated risks. Results: The ability of LLMs to reproduce human expert consensus varied widely with some models demonstrating a balanced approach to risk in a similar manner to humans, whilst other models reported outputs at either end of the spectrum for risk tolerance. In terms of agreement with the human mode, Chat GPT 3.5 and GPT 4.0 (Open AI) had the highest levels of alignment, with the human mode selected in 6/11 questions. The total biopsy propensity score generated from the human mode was 23/44. Both Open AI models produced similar propensity scores between 22 and 24, however Llama 2 and MS Co-pilot also reported scores within this range, but with poorer response alignment to the human mode at only 2/11 questions. The most risk averse model in this study was MedLM with a propensity score of 11 and the least risk averse model was Claude 3 with a score of 34. Conclusions: LLM outputs demonstrated a modest ability to replicate human clinical decision making in this study, however the performance varied widely between LLM models. Questions with more uniform human responses produced LLM outputs with greater alignment, whereas in questions with low levels of human consensus there was poor output alignment. This may limit the practical use of LLMs in real world clinical practice.

Navigating the Boundaries of Teleconsultation—Capabilities, Limitations, and Pathways for Improvement: Qualitative Study of the Experiences of Patients With Stroke

Background: Survivors of stroke often face persistent challenges accessing postdischarge care due to mobility limitations, transportation burdens, and inflexible scheduling. Teleconsultation has emerged as a potential solution to improve continuity of care, but its perceived strengths and limitations from the patient perspective remain insufficiently understood. Objective: This study aimed to explore the experiences of survivors of stroke with a nurse-led teleconsultation program to (1) identify perceived capabilities; (2) understand limitations in usability, accessibility, and clinical function; and (3) generate patient-informed recommendations for improvement. Methods: A qualitative study was embedded within a 3-month nurse-led teleconsultation intervention delivered by advanced practice nurses. A total of 21 survivors of ischemic stroke (aged 45-76 y; female: n=11, 52%) who had preserved cognitive function (Montreal Cognitive Assessment score ≥22) and smartphone access participated in 6 focus groups conducted via Zoom. Data were analyzed thematically using an established framework. Data saturation was achieved. Results: Participants widely valued teleconsultation for reducing logistical burdens; enhancing access; and offering a more comfortable, emotionally supportive setting for follow-up care. Many reported increased awareness and motivation for self-monitoring. However, limitations included an inability to perform physical assessments or respond to emergencies; digital and usability barriers, especially among older users; and scheduling inflexibility. Participants emphasized the need for patient-initiated follow-up mechanisms, physician collaboration for medication management, and greater support for users considered digitally marginalized. They also highlighted the potential of teleconsultation to serve as a triage tool, reserving in-person care for complex cases. Conclusions: Nurse-led teleconsultation was perceived as a convenient and supportive modality for poststroke care, particularly for stable follow-ups and psychosocial support. However, its long-term viability depends on addressing clinical and technical limitations, enhancing user autonomy, and integrating interdisciplinary input. By centering the lived experiences of survivors of stroke, this study offers concrete recommendations to guide the development of more inclusive, responsive, and patient-centered teleconsultation models.

Developing an Evaluation System for Quality of Health Educational Short Videos on Social Media (LassVQ) Using Nominal Group Technique and Analytic Hierarchy Process: Qualitative Study

Background: With the increasing use of social media platforms for health communication, the quality of health educational short videos (HESVs) has become a key concern. However, no standardized framework exists to evaluate the quality of health videos on social media, highlighting the need for a comprehensive evaluation system. Objective: The aim of this study is to develop a valid and structured evaluation tool for assessing the quality of HESVs on social media. Methods: The initial evaluation indicators obtained from the literature review and brainstorming undertaken in the study group were provided to the nominal group reference Lasswell’s 5W communication model, and two rounds of nominal group technique (NGT) were carried out to screen, add, revise, and adjust indicators, and reach a consensus of evaluation system. The indicators were then ranked based on their significance, as scored by the experts using the analytic hierarchy process. The content validity was assessed by experts who rated the relevance of each indicator on a 4-point Likert scale. Results: The primary indicators include communicator, communication content, communication channel, and communication effect, along with 13 secondary indicators and 34 tertiary indicators. 11 experts were enrolled in the NGT, 45% of experts had a doctoral degree, 80% of them were ranked associate professor or professor. The average familiarity coefficient of each key indicator of the NGT was 0.85. The average values of the expert judgment coefficient and authority coefficient were 0.93 and 0.85, respectively. In Round 1 of NGT, the “Communication target” of 5 primary indicators, 7 of 20 secondary indicators, and 66 of 94 tertiary indicators did not reach a consensus, and therefore, they were not deleted and will proceed to the next round of NGT. In Round 2 NGT, 1 primary indicator, 7 secondary indicators, and 59 tertiary indicators were deleted based on the consensus criteria. After the two rounds of NGT, 4 primary indicators, 13 secondary indicators, and 34 tertiary indicators finally reached a consensus. Among primary indicators, communication content was found to be the most influential, accounting for 45.68%. Among secondary indicators, credibility, scientificity, availability, and social attention were the most influential indicators, with priorities of 56.67%, 24.26%, 74.62%, and 39.89% in their respective categories. Among tertiary indicators, ‘Become a hot search recommended by the platform’ was the most influential indicator with a weight of 0.07. The content validity of all the evaluation indicators were 0.73 – 1.0, and the scale-level content validity index (average) was 0.87, which was indicated as acceptable. Conclusions: The evaluation system for the quality of HESVs on social media (LassVQ) was developed, and its validity was acceptable. The proposed evaluation system can be used in conjunction with qualitative methods to gain a holistic perspective on the multidimensional quality of HESVs on social media.
  • ✇Cell
  • Multi-adjuvant personalized neoantigen vaccines: Fine-tuning anti-cancer T cells Hejia Henry Wang · Neeha Zaidi
    Personalized cancer vaccines aim to broaden the anti-tumor T cell repertoire by targeting neoantigens unique to each patient’s tumor, but immunogenicity has been inconsistent. In this issue of Cell, Blass, Keskin, Tu et al. evaluate NeoVaxMI, a multi-adjuvant personalized synthetic long-peptide vaccine administered with nivolumab in patients with melanoma. NeoVaxMI elicited stronger CD4+ and CD8+ responses than earlier iterations, and vaccine-induced T cells trafficked to regressing metastatic l
     

Multi-adjuvant personalized neoantigen vaccines: Fine-tuning anti-cancer T cells

18 September 2025 at 08:00
Personalized cancer vaccines aim to broaden the anti-tumor T cell repertoire by targeting neoantigens unique to each patient’s tumor, but immunogenicity has been inconsistent. In this issue of Cell, Blass, Keskin, Tu et al. evaluate NeoVaxMI, a multi-adjuvant personalized synthetic long-peptide vaccine administered with nivolumab in patients with melanoma. NeoVaxMI elicited stronger CD4+ and CD8+ responses than earlier iterations, and vaccine-induced T cells trafficked to regressing metastatic lesions.

The arts for disease prevention and health promotion: a systematic review

Nature Medicine, Published online: 18 September 2025; doi:10.1038/s41591-025-03962-7

The arts, according to a systematic synthesis of data from 95 studies (across 26 countries), may support non-communicable disease prevention by providing opportunities for increased physical activity, and helping to address social forces that contribute to health inequities.
❌