❌

Normal view

VaultGemma: A Differentially Private Gemma Model

arXiv:2510.15001v2 Announce Type: replace-cross Abstract: We introduce VaultGemma 1B, a 1 billion parameter model within the Gemma family, fully trained with differential privacy. Pretrained on the identical data mixture used for the Gemma 2 series, VaultGemma 1B represents a significant step forward in privacy-preserving large language models. We openly release this model to the community

Framework for the Development and Delivery of Digital Peer Support Programs: Qualitative Study on in-Person and Digital Delivery for People With Cardiovascular Disease

Background: Peer support (sharing experiences/support with others with the same condition) improves health outcomes among people with cardiovascular disease (CVD), including self-management behaviours and self-efficacy. However, current peer support interventions are diverse. Evidence is lacking on peer support attenders perceptions of benefits and the elements that are considered priorities, especially for digital interventions. Objective: The study objectives were to 1) describe perceived benefits and recommendations for CVD peer support programs from people attending in-person peer support, 2) identify priorities for digital peer support from consumers and clinicians testing a peer support app prototype, and 3) develop a framework to inform future peer support intervention development. Methods: Qualitative methodology was used across two components to address the objectives of this study. In Component 1, semi-structured focus groups were conducted with attenders of established in-person CVD peer support groups, exploring the perceived benefits of peer support and recommendations for future programs. In Component 2, semi-structured interactive workshops with consumers with CVD and semi-structured online interviews with CVD clinicians/researchers were undertaken seeking feedback and recommendations for digital peer support using an exploratory digital CVD peer support application prototype. Data were recorded digitally, transcribed verbatim, and analysed thematically. Findings from both components were iteratively synthesised to inform a digital peer support development framework. Results: In Component 1, 22 participants (age range 29-84 years, male 45%) took part in focus groups. The overarching theme was that peer support provides benefits through sharing experiences. Five themes were refined and defined; (i) peer support provides a way of coping, (ii) peers learn from each other, (iii) peers understand what each other are going through, (iv) the peer community uplifts mood and build confidence, and (v) awareness, flexibility and resources are important for engagement. In Component 2, five participants (age range 55-74 years, male 60%) attended two workshops and eight clinicians/researchers (age range 30-65 years, male 10%) were interviewed. Three themes were refined and defined: (i) autonomy is essential to promote engagement, (ii) safeguarding is important to both users and clinicians, and (iii) interfaces that are simple, easy to use and visually attractive enable use. Priorities identified from both components included greater peer support awareness and uptake, flexibility with timing and family participation, healthcare professional involvement, provision of resources, autonomous features enabling choice, checklists and clinician moderation for safeguarding, and simple to use interfaces. Conclusions: Participants in peer support programs derive benefit from sharing their experience of living with CVD which enable coping, learning, feeling understood and a sense of community. Priorities were synthesised to create a framework for digital peer support development for future peer support with recommendations to focus on six key areas: uptake, flexibility, resources, autonomy, safeguarding and interface.

Efficient and accurate search in petabase-scale sequence repositories

Nature, Published online: 08 October 2025; doi:10.1038/s41586-025-09603-w

MetaGraph enables scalable indexing of large sets of DNA, RNA or protein sequences using annotated de Bruijn graphs.

Pathobiology and Genetics

Pneumologie. 2025 Oct;79(10):701-711. doi: 10.1055/a-2625-4648. Epub 2025 Oct 6.

ABSTRACT

Genetics and pathobiology were addressed at the 7th World Symposium on Pulmonary Hypertension in Task Forces 2 and 3. The Genetics Task Force also focused on precision medicine approaches, and the Pathobiology working group concentrated heavily on new omics technologies. Therefore, the following not only summarises the current state of knowledge on genetics, genetic testing methods, and molecular pathophysiological changes, but also places it in context and critically discusses it. In addition, the importance of national and international biobanks and cohorts, as well as the active involvement of patients and families, is emphasized.

PMID:41052524 | DOI:10.1055/a-2625-4648

Evaluating Large Language Models and Retrieval-Augmented Generation Enhancement for Delivering Guideline-Adherent Nutrition Information for Cardiovascular Disease Prevention: Cross-Sectional Study

Background: Cardiovascular disease (CVD) remains the leading cause of death worldwide, yet many web-based sources on cardiovascular (CV) health are inaccessible. Large language models (LLMs) are increasingly used for health-related inquiries and offer an opportunity to produce accessible and scalable CV health information. However, because these models are trained on heterogeneous data, including unverified user-generated content, the quality and reliability of food and nutrition information on CVD prevention remain uncertain. Recent studies have examined LLM use in various health care applications, but their effectiveness for providing nutrition information remains understudied. Although retrieval-augmented generation (RAG) frameworks have been shown to enhance LLM consistency and accuracy, their use in delivering nutrition information for CVD prevention requires further evaluation. Objective: To evaluate the effectiveness of off-the-shelf and RAG-enhanced LLMs in delivering guideline-adherent nutrition information for CVD prevention, we assessed 3 off-the-shelf models (ChatGPT-4o, Perplexity, and Llama 3-70B) and a Llama 3-70B+RAG model. Methods: We curated 30 nutrition questions that comprehensively addressed CVD prevention. These were approved by a registered dietitian providing preventive cardiology services at an academic medical center and were posed 3 times to each model. We developed a 15,074-word knowledge bank incorporating the American Heart Association’s 2021 dietary guidelines and related website content to enhance Meta’s Llama 3-70B model using RAG. The model received this and a few-shot prompt as context, included citations in a Context Source section, and used vector similarity to align responses with guideline content, with the temperature parameter set to 0.5 to enhance consistency. Model responses were evaluated by 3 expert reviewers against benchmark CV guidelines for appropriateness, reliability, readability, harm, and guideline adherence. Mean scores were compared using ANOVA, with statistical significance set at P<.05. interrater agreement was measured using the cohen coefficient and readability estimated flesch-kincaid score. results: llama model scored higher than perplexity gpt-4o models on reliability appropriateness guideline adherence showed no harm.>70%; P<.001 indicated high reviewer agreement. conclusions: the llama model outperformed off-the-shelf models across all measures with no evidence of harm although responses were less readable due to technical language. scored lower on and produced some harmful responses. these findings highlight limitations demonstrate that rag system integration can enhance llm performance in delivering evidence-based dietary information.>

The Role of Data in Public Health and Health Innovation: Perspectives on Social Determinants of Health, Community-Based Data Approaches, and AI

Public health is undergoing profound transformation driven by data from the global health sector and related fields. To address systemic health disparities, scholars and practitioners are increasingly applying a data equity lens, an approach that has become even more urgent as the United States faces the erosion of public health data infrastructure. This paper summarizes insights from an April 2024 convening by the Yale School of Public Health—The Role of Data in Public Health Equity and Innovation—with intersectoral stakeholders from academia, government (local, state, and federal), healthcare, and private industry. The convening included keynote presentations and roundtables regarding the depiction of social determinants of health (SDOH) in data; effects of artificial intelligence (AI) on health data equity; and community-based models for data, providing a framework for cross-cutting discussions. Through a narrative synthesis, themes were identified and synthesized from systematically gathered information from presentations and roundtables. This process led to a set of actionable, cross-cutting recommendations to guide inclusive and impactful data practices for policymakers, public health professionals, and health innovators across diverse contexts: (1) Enable big data and interoperability connecting SDOH and health outcomes; (2) Include diverse, non-technical voices in AI and health discussions; (3) Fund research on data equity and AI in health sciences; (4) Modernize Health Insurance Portability and Accountability Act (HIPAA) with new guidelines for AI and big data; and (5) Research and conceptual frameworks are needed to elucidate interconnections between data equity and health equity.

Generative artificial intelligence in medicine

Nature Medicine, Published online: 06 October 2025; doi:10.1038/s41591-025-03983-2

This Review summarizes recent technical advancements in generative AI, outlines how new models might improve healthcare and discusses validation approaches—using lessons from recent successes and failures in the field.

Exploring Attitudes and Obstacles Around Digital Public Health Tools: Insights From a Statewide Cross-Sectional Survey on Washington’s Vaccine Verification System

Background: Development and use of digital public health tools surged during the COVID-19 pandemic. Among these tools, vaccine verification systems emerged as alternatives to paper vaccine records, aiming to help limit the spread of disease. In November 2021, the Washington State Department of Health launched “WA Verify,” a QR code–based vaccine verification system built on the SMART Health Card framework, providing residents with a convenient way to store and share proof of vaccination digitally. However, WA Verify was developed and deployed before assessments and public input regarding potential adoption challenges—such as concerns about privacy, surveillance, data sharing, trust in the technology, and the managing organizations—could be completed. Objective: This analysis used statewide survey data from Washington to identify and characterize barriers and facilitators to the adoption of WA Verify, and to understand how factors such as data privacy, security, attitudes toward public health policies and communication, and technological proficiency may influence acceptance and uptake of digital public health tools. Methods: A cross-sectional statewide survey was distributed between September 2022 and January 2023 to a random sample of 5000 Washington households. Respondents were categorized into 3 groups based on their responses indicating WA Verify “users,” “potential users,” or “unlikely users.” Comparisons were made between groups regarding experiences with and opinions on COVID-19 vaccine and test verification, public health policies, communication, digital tools, technological proficiency, sociodemographic characteristics, and health history. Poststratification weights were applied to reduce nonresponse bias. Results: Of the 1401 respondents, 359 (25.6% unweighted, 25.8% weighted) were users, 662 (47.3% unweighted, 49.8% weighted) were potential users, and 380 (27.1% unweighted, 24.4% weighted) were unlikely users. All percentages reported are based on weighted data. Compared with users and potential users, unlikely users were more likely to oppose policies requiring proof of COVID-19 vaccination or negative test results (users: 6.0%, potential users: 13.6%, unlikely users: 65.9%). Unlikely users were more likely to cite concerns about personal health data security and phone hacking or tracking, though these concerns were also notable among potential users and users. Users and potential users were more likely to perceive a digital vaccine verification system as convenient (users: 96.5%, potential users: 92.3%, unlikely users: 38.1%) and indicated openness to receiving relevant information from a range of sources. Unlikely users were more likely to report not owning a smartphone and demonstrated lower technological proficiency (users: 12.3%, potential users: 15.9%, unlikely users: 32.3%), indicating a technological divide between groups. Conclusions: While nearly three-quarters of respondents had either already adopted or were willing to adopt a tool like WA Verify, concerns about data security, lower technological proficiency, and distrust of public health characterized those least likely to adopt such tools. Identifying barriers to adoption among “unlikely users” is essential for developing effective communication strategies—such as targeted marketing and community engagement—to improve adoption and ensure equitable access to public health technologies.

Comparative Evaluation of a Medical Large Language Model in Answering Real-World Radiation Oncology Questions: Multicenter Observational Study

Background: Large language models (LLMs) hold promise for supporting clinical tasks, particularly in data-driven and technical disciplines such as radiation oncology. While prior evaluation studies have focused on examination-style settings for evaluating LLMs, their performance in real-life clinical scenarios remains unclear. In the future, LLMs might be used as general AI assistants to answer questions arising in clinical practice. It is unclear how well a modern LLM, locally executed within the infrastructure of a hospital, would answer such questions compared with clinical experts. Objective: This study aimed to assess the performance of a locally deployed, state-of-the-art medical LLM in answering real-world clinical questions in radiation oncology compared with clinical experts. The aim was to evaluate the overall quality of answers, as well as the potential harmfulness of the answers if used for clinical decision-making. Methods: Physicians from 10 departments of European hospitals collected questions arising in the clinical practice of radiation oncology. Fifty of these questions were answered by 3 senior radiation oncology experts with at least 10 years of work experience, as well as the LLM Llama3-OpenBioLLM-70B (Ankit Pal and Malaikannan Sankarasubbu). In a blinded review, physicians rated the overall answer quality on a 5-point Likert scale (quality), assessed whether an answer might be potentially harmful if used for clinical decision-making (harmfulness), and determined if responses were from an expert or the LLM (recognizability). Comparisons between clinical experts and LLMs were then made for quality, harmfulness, and recognizability. Results: There were no significant differences between the quality of the answers between LLM and clinical experts (mean scores of 3.38 vs 3.63; median 4.00, IQR 3.00-4.00 vs median 3.67, IQR 3.33-4.00; P=.26; Wilcoxon signed rank test). The answers were deemed potentially harmful in 13% of cases for the clinical experts compared with 16% of cases for the LLM (P=.63; Fisher exact test). Physicians correctly identified whether an answer was given by a clinical expert or an LLM in 78% and 72% of cases, respectively. Conclusions: A state-of-the-art medical LLM can answer real-life questions from the clinical practice of radiation oncology similarly well as clinical experts regarding overall quality and potential harmfulness. Such LLMs can already be deployed within the local hospital environment at an affordable cost. While LLMs may not yet be ready for clinical implementation as general AI assistants, the technology continues to improve at a rapid pace. Evaluation studies based on real-life situations are important to better understand the weaknesses and limitations of LLMs in clinical practice. Such studies are also crucial to define when the technology is ready for clinical implementation. Furthermore, education for health care professionals on generative AI is needed to ensure responsible clinical implementation of this transforming technology.

Large Language Models’ Clinical Decision-Making on When to Perform a Kidney Biopsy: Comparative Study

Background: Artificial intelligence (AI) and Large Language models (LLMs) are increasing in sophistication and are being integrated into many disciplines. The potential for LLMs to augment clinical decisions is an evolving area of research. Objective: This study compared the responses of over 1000 kidney specialist physicians (nephrologists) to outputs of commonly used LLMs using a questionnaire determining when a kidney biopsy should be performed. Methods: This research group completed a large online questionnaire for nephrologists to determine when a kidney biopsy should be performed. The questionnaire was co-designed with patient participation, refined through multiple iterations, then piloted locally before international dissemination. It was the largest international study in the field and demonstrated variation between human clinicians in biopsy propensity relating to human factors such as sex and age, as well as systemic factors such as country, job seniority and technical proficiency. The same questions were put to both human doctors and LLMs in an identical order in a single session. Eight commonly used LLMs were interrogated: Chat GPT 3.5, Mistral Hugging Face, Perplexity, Microsoft Co-pilot, Llama 2, GPT 4.0, MedLM and Claude 3. The most common response given by clinicians (human mode) to each question was taken as the baseline for comparison. Questionnaire responses to the indications and contraindications for biopsy generated a score (0-44) reflecting biopsy propensity, in which a higher score was used as a surrogate marker for an increased tolerance of potential associated risks. Results: The ability of LLMs to reproduce human expert consensus varied widely with some models demonstrating a balanced approach to risk in a similar manner to humans, whilst other models reported outputs at either end of the spectrum for risk tolerance. In terms of agreement with the human mode, Chat GPT 3.5 and GPT 4.0 (Open AI) had the highest levels of alignment, with the human mode selected in 6/11 questions. The total biopsy propensity score generated from the human mode was 23/44. Both Open AI models produced similar propensity scores between 22 and 24, however Llama 2 and MS Co-pilot also reported scores within this range, but with poorer response alignment to the human mode at only 2/11 questions. The most risk averse model in this study was MedLM with a propensity score of 11 and the least risk averse model was Claude 3 with a score of 34. Conclusions: LLM outputs demonstrated a modest ability to replicate human clinical decision making in this study, however the performance varied widely between LLM models. Questions with more uniform human responses produced LLM outputs with greater alignment, whereas in questions with low levels of human consensus there was poor output alignment. This may limit the practical use of LLMs in real world clinical practice.

New Doc on the Block: Scoping Review of AI Systems Delivering Motivational Interviewing for Health Behavior Change

Background: Artificial intelligence (AI) is increasingly used in digital health, particularly through large language models (LLMs), to support patient engagement and behavior change. One novel application is the delivery of motivational interviewing (MI), an evidence-based, patient-centered counseling technique designed to enhance motivation and resolve ambivalence around health behaviors. AI tools, including chatbots, mobile apps, and web-based agents, are being developed to simulate MI techniques at scale. While these innovations are promising, important questions remain about how faithfully AI systems can replicate MI principles or achieve meaningful behavioral impact. Objective: This scoping review aimed to summarize existing empirical studies evaluating AI-driven systems that apply MI techniques to support health behavior change. Specifically, we examined the feasibility of these systems; their fidelity to MI principles; and their reported behavioral, psychological, or engagement outcomes. Methods: We systematically searched PubMed, Embase, Scopus, Web of Science, and Cochrane Library for empirical studies published between January 1, 2018, and February 25, 2025. Eligible studies involved AI-driven systems using natural language generation, understanding, or computational logic to deliver MI techniques to users targeting a specific health behavior. We excluded studies using AI solely for training clinicians in MI. Three independent reviewers screened and extracted data on study design, AI modality and type, MI components, health behavior focus, MI fidelity assessment, and outcome domains. Results: Of the 1001 records identified, 15 (1.5%) met the inclusion criteria. Of these 15 studies, 6 (40%) were exploratory feasibility or pilot studies, and 3 (20%) were randomized controlled trials. AI modalities included rule-based chatbots (9/15, 60%), LLM-based systems (4/15, 27%), and virtual or mobile agents (2/15, 13%). Targeted behaviors included smoking cessation (6/15, 40%), substance use (3/15, 20%), COVID-19 vaccine hesitancy, type 2 diabetes self-management, stress, mental health service use, and opioid use during pregnancy. Of the 15 studies, 13 (87%) reported positive findings on feasibility or user acceptability, while 6 (40%) assessed MI fidelity using expert review or structured coding, with moderate to high alignment reported. Several studies found that users perceived the AI systems as judgment free, supportive, and easier to engage with than human counselors, particularly in stigmatized contexts. However, limitations in empathy, safety transparency, and emotional nuance were commonly noted. Only 3 (20%) of the 15 studies reported substantially significant behavioral changes. Conclusions: AI systems delivering MI show promise for enhancing patient engagement and scaling behavior change interventions. Early evidence supports their usability and partial fidelity to MI principles, especially in sensitive domains. However, most systems remain in early development, and few have been rigorously tested. Future research should prioritize randomized evaluations; standardized fidelity measures; and safeguards for LLM safety, empathy, and accuracy in health-related dialogue. Trial Registration: OSF Registries 10.17605/OSF.IO/G9N7E; https://osf.io/g9n7e

GASPS: A Multi-Omics Framework for Defining Genomic Aberration-Driven Signatures and Predicting Patient Outcomes in Lung Cancer

bioRxiv [Preprint]. 2025 Aug 25:2025.08.21.671519. doi: 10.1101/2025.08.21.671519.

ABSTRACT

Lung cancer is the most common cause of cancer-related death worldwide. Recent advancements in targeted therapies and immunotherapies have achieved remarkable success. However, patient responses to treatments with lung cancer vary substantially. The mutation status of driver genes can direct personalized treatment, but their prognostic value and treatment efficacy are limited. In this study, we developed a statistical framework named Genomic Aberration-Derived Signature for Patient Stratification (GASPS) to characterize the transcriptomic deregulation of driver genomic aberrations and stratify patients. By applying GASPS to The Cancer Genome Atlas Lung Adenocarcinoma (TCGA-LUAD) data, we developed gene signatures for 38 driver genomic aberrations, including gene mutations, amplifications, and deletions. These signatures were applied to independent lung cancer transcriptomic datasets containing a total of 2,226 patient samples. Our results indicated that these driver gene signatures are much more prognostic than their corresponding genomic mutations. Interestingly, the two EGFR-related signatures characterizing EGFR mutation and amplification, respectively, exhibited contrasting associations with prognosis, treatment response, and immune infiltration in the tumor microenvironment. Moreover, the STK11 mutation signature, rather than the mutation status, was found to be predictive of the response and long-term benefit of patients treated with immune checkpoint blockade therapy in lung cancer. This framework is readily applicable to most cancer types using existing data to improve prognostic risk assessment and treatment efficacy by guiding personalized therapies.

PMID:40909579 | PMC:PMC12407784 | DOI:10.1101/2025.08.21.671519

GASPS: A Multi-Omics Framework for Defining Genomic Aberration-Driven Signatures and Predicting Patient Outcomes in Lung Cancer

bioRxiv [Preprint]. 2025 Aug 25:2025.08.21.671519. doi: 10.1101/2025.08.21.671519.

ABSTRACT

Lung cancer is the most common cause of cancer-related death worldwide. Recent advancements in targeted therapies and immunotherapies have achieved remarkable success. However, patient responses to treatments with lung cancer vary substantially. The mutation status of driver genes can direct personalized treatment, but their prognostic value and treatment efficacy are limited. In this study, we developed a statistical framework named Genomic Aberration-Derived Signature for Patient Stratification (GASPS) to characterize the transcriptomic deregulation of driver genomic aberrations and stratify patients. By applying GASPS to The Cancer Genome Atlas Lung Adenocarcinoma (TCGA-LUAD) data, we developed gene signatures for 38 driver genomic aberrations, including gene mutations, amplifications, and deletions. These signatures were applied to independent lung cancer transcriptomic datasets containing a total of 2,226 patient samples. Our results indicated that these driver gene signatures are much more prognostic than their corresponding genomic mutations. Interestingly, the two EGFR-related signatures characterizing EGFR mutation and amplification, respectively, exhibited contrasting associations with prognosis, treatment response, and immune infiltration in the tumor microenvironment. Moreover, the STK11 mutation signature, rather than the mutation status, was found to be predictive of the response and long-term benefit of patients treated with immune checkpoint blockade therapy in lung cancer. This framework is readily applicable to most cancer types using existing data to improve prognostic risk assessment and treatment efficacy by guiding personalized therapies.

PMID:40909579 | PMC:PMC12407784 | DOI:10.1101/2025.08.21.671519

Scalable generation and functional classification of genetic variants in inborn errors of immunity to accelerate clinical diagnosis and treatment

In lieu of traditional genetic variant testing approaches, an approach using scalable variant classification in primary human T cells with a clinically relevant readout can inform rapid diagnosis and treatment of inborn errors of immunity.
❌