❌

Normal view

The Quest for Reliable Metrics of Responsible AI

arXiv:2510.26007v1 Announce Type: cross Abstract: The development of Artificial Intelligence (AI), including AI in Science (AIS), should be done following the principles of responsible AI. Progress in responsible AI is often quantified through evaluation metrics, yet there has been less work on assessing the robustness and reliability of the metrics themselves. We reflect on prior work that examines the robustness of fairness metrics for recommender systems as a type of AI application and summarise their key takeaways into a set of non-exhaustive guidelines for developing reliable metrics of responsible AI. Our guidelines apply to a broad spectrum of AI applications, including AIS.

Multi-omic profiling reveals age-related immune dynamics in healthy adults

Nature, Published online: 29 October 2025; doi:10.1038/s41586-025-09686-5

This multi-omic longitudinal analysis of the healthy human peripheral immune system constructs the Human Immune Health Atlas and assembles data on immune cell composition and state changes with age, including responses to cytomegalovirus infection and influenza vaccination.

Dynamic Monitoring of Recurrent Ovarian Cancer Using Serial ctDNA: A Real-World Case Series

Curr Oncol. 2025 Oct 21;32(10):585. doi: 10.3390/curroncol32100585.

ABSTRACT

Recurrent ovarian cancer (OC) is challenging to detect early using current methods like CA-125 and imaging. Circulating tumor DNA (ctDNA) may improve disease monitoring. Here, we assess the real-world clinical utility of serial ctDNA analyses in patients with recurrent OC. We analyzed serial plasma samples (N = 23) from six patients with recurrent OC using a tumor-informed next-generation sequencing assay targeting 68 cancer-related genes developed at the University of Washington. ctDNA variant allele frequencies (VAFs) were correlated with CA-125 levels, radiographic findings, and clinical outcomes. ctDNA levels generally reflected clinical status, accurately mirroring disease progression and therapeutic response. In one patient, rising ctDNA preceded clinical recurrence by four months, despite normal CA-125 and imaging, highlighting its potential advantage. Conversely, some patients exhibited clinical progression with undetectable ctDNA, indicating limitations in assay sensitivity, biological factors, or metastatic sites (e.g., brain metastases). ctDNA and CA-125 showed complementary value in most cases, suggesting potential combined use in clinical monitoring. Our findings demonstrate that ctDNA is a promising biomarker to complement existing monitoring approaches for recurrent OC. In some cases, capable of predicting relapse and treatment response ahead of current clinical indicators. However, identified discordances underscore technical and biological challenges that warrant further investigation. Larger prospective studies are necessary to refine ctDNA's clinical utility and integration into personalized OC care.

PMID:41149505 | PMC:PMC12563156 | DOI:10.3390/curroncol32100585

Reduced AI Acceptance After the Generative AI Boom: Evidence From a Two-Wave Survey Study

arXiv:2510.23578v1 Announce Type: new Abstract: The rapid adoption of generative artificial intelligence (GenAI) technologies has led many organizations to integrate AI into their products and services, often without considering user preferences. Yet, public attitudes toward AI use, especially in impactful decision-making scenarios, are underexplored. Using a large-scale two-wave survey study (n_wave1=1514, n_wave2=1488) representative of the Swiss population, we examine shifts in public attitudes toward AI before and after the launch of ChatGPT. We find that the GenAI boom is significantly associated with reduced public acceptance of AI (see Figure 1) and increased demand for human oversight in various decision-making contexts. The proportion of respondents finding AI "not acceptable at all" increased from 23% to 30%, while support for human-only decision-making rose from 18% to 26%. These shifts have amplified existing social inequalities in terms of widened educational, linguistic, and gender gaps post-boom. Our findings challenge industry assumptions about public readiness for AI deployment and highlight the critical importance of aligning technological development with evolving public preferences.

ReXGroundingCT: A 3D Chest CT Dataset for Segmentation of Findings from Free-Text Reports

arXiv:2507.22030v2 Announce Type: replace-cross Abstract: We introduce ReXGroundingCT, the first publicly available dataset linking free-text findings to pixel-level 3D segmentations in chest CT scans. The dataset includes 3,142 non-contrast chest CT scans paired with standardized radiology reports from CT-RATE. Construction followed a structured three-stage pipeline. First, GPT-4 was used to extract and standardize findings, descriptors, and metadata from reports originally written in Turkish and machine-translated into English. Second, GPT-4o-mini categorized each finding into a hierarchical ontology of lung and pleural abnormalities. Third, 3D annotations were produced for all CT volumes: the training set was quality-assured by board-certified radiologists, and the validation and test sets were fully annotated by board-certified radiologists. Additionally, a complementary chain-of-thought dataset was created to provide step-by-step hierarchical anatomical reasoning for localizing findings within the CT volume, using GPT-4o and localization coordinates derived from organ segmentation models. ReXGroundingCT contains 16,301 annotated entities across 8,028 text-to-3D-segmentation pairs, covering diverse radiological patterns from 3,142 non-contrast CT scans. About 79% of findings are focal abnormalities and 21% are non-focal. The dataset includes a public validation set of 50 cases and a private test set of 100 cases, both annotated by board-certified radiologists. The dataset establishes a foundation for enabling free-text finding segmentation and grounded radiology report generation in CT imaging. Model performance on the private test set is hosted on a public leaderboard at https://rexrank.ai/ReXGroundingCT. The dataset is available at https://huggingface.co/datasets/rajpurkarlab/ReXGroundingCT.

Benchmarking large language models for personalized, biomarker-based health intervention recommendations

npj Digital Medicine, Published online: 27 October 2025; doi:10.1038/s41746-025-01996-2

Benchmarking large language models for personalized, biomarker-based health intervention recommendations

A Definition of AGI

arXiv:2510.18212v2 Announce Type: replace Abstract: The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quantifiable framework to address this, defining AGI as matching the cognitive versatility and proficiency of a well-educated adult. To operationalize this, we ground our methodology in Cattell-Horn-Carroll theory, the most empirically validated model of human cognition. The framework dissects general intelligence into ten core cognitive domains-including reasoning, memory, and perception-and adapts established human psychometric batteries to evaluate AI systems. Application of this framework reveals a highly "jagged" cognitive profile in contemporary models. While proficient in knowledge-intensive domains, current AI systems have critical deficits in foundational cognitive machinery, particularly long-term memory storage. The resulting AGI scores (e.g., GPT-4 at 27%, GPT-5 at 57%) concretely quantify both rapid progress and the substantial gap remaining before AGI.

VaultGemma: A Differentially Private Gemma Model

arXiv:2510.15001v2 Announce Type: replace-cross Abstract: We introduce VaultGemma 1B, a 1 billion parameter model within the Gemma family, fully trained with differential privacy. Pretrained on the identical data mixture used for the Gemma 2 series, VaultGemma 1B represents a significant step forward in privacy-preserving large language models. We openly release this model to the community

Framework for the Development and Delivery of Digital Peer Support Programs: Qualitative Study on in-Person and Digital Delivery for People With Cardiovascular Disease

Background: Peer support (sharing experiences/support with others with the same condition) improves health outcomes among people with cardiovascular disease (CVD), including self-management behaviours and self-efficacy. However, current peer support interventions are diverse. Evidence is lacking on peer support attenders perceptions of benefits and the elements that are considered priorities, especially for digital interventions. Objective: The study objectives were to 1) describe perceived benefits and recommendations for CVD peer support programs from people attending in-person peer support, 2) identify priorities for digital peer support from consumers and clinicians testing a peer support app prototype, and 3) develop a framework to inform future peer support intervention development. Methods: Qualitative methodology was used across two components to address the objectives of this study. In Component 1, semi-structured focus groups were conducted with attenders of established in-person CVD peer support groups, exploring the perceived benefits of peer support and recommendations for future programs. In Component 2, semi-structured interactive workshops with consumers with CVD and semi-structured online interviews with CVD clinicians/researchers were undertaken seeking feedback and recommendations for digital peer support using an exploratory digital CVD peer support application prototype. Data were recorded digitally, transcribed verbatim, and analysed thematically. Findings from both components were iteratively synthesised to inform a digital peer support development framework. Results: In Component 1, 22 participants (age range 29-84 years, male 45%) took part in focus groups. The overarching theme was that peer support provides benefits through sharing experiences. Five themes were refined and defined; (i) peer support provides a way of coping, (ii) peers learn from each other, (iii) peers understand what each other are going through, (iv) the peer community uplifts mood and build confidence, and (v) awareness, flexibility and resources are important for engagement. In Component 2, five participants (age range 55-74 years, male 60%) attended two workshops and eight clinicians/researchers (age range 30-65 years, male 10%) were interviewed. Three themes were refined and defined: (i) autonomy is essential to promote engagement, (ii) safeguarding is important to both users and clinicians, and (iii) interfaces that are simple, easy to use and visually attractive enable use. Priorities identified from both components included greater peer support awareness and uptake, flexibility with timing and family participation, healthcare professional involvement, provision of resources, autonomous features enabling choice, checklists and clinician moderation for safeguarding, and simple to use interfaces. Conclusions: Participants in peer support programs derive benefit from sharing their experience of living with CVD which enable coping, learning, feeling understood and a sense of community. Priorities were synthesised to create a framework for digital peer support development for future peer support with recommendations to focus on six key areas: uptake, flexibility, resources, autonomy, safeguarding and interface.

Efficient and accurate search in petabase-scale sequence repositories

Nature, Published online: 08 October 2025; doi:10.1038/s41586-025-09603-w

MetaGraph enables scalable indexing of large sets of DNA, RNA or protein sequences using annotated de Bruijn graphs.

Pathobiology and Genetics

Pneumologie. 2025 Oct;79(10):701-711. doi: 10.1055/a-2625-4648. Epub 2025 Oct 6.

ABSTRACT

Genetics and pathobiology were addressed at the 7th World Symposium on Pulmonary Hypertension in Task Forces 2 and 3. The Genetics Task Force also focused on precision medicine approaches, and the Pathobiology working group concentrated heavily on new omics technologies. Therefore, the following not only summarises the current state of knowledge on genetics, genetic testing methods, and molecular pathophysiological changes, but also places it in context and critically discusses it. In addition, the importance of national and international biobanks and cohorts, as well as the active involvement of patients and families, is emphasized.

PMID:41052524 | DOI:10.1055/a-2625-4648

Evaluating Large Language Models and Retrieval-Augmented Generation Enhancement for Delivering Guideline-Adherent Nutrition Information for Cardiovascular Disease Prevention: Cross-Sectional Study

Background: Cardiovascular disease (CVD) remains the leading cause of death worldwide, yet many web-based sources on cardiovascular (CV) health are inaccessible. Large language models (LLMs) are increasingly used for health-related inquiries and offer an opportunity to produce accessible and scalable CV health information. However, because these models are trained on heterogeneous data, including unverified user-generated content, the quality and reliability of food and nutrition information on CVD prevention remain uncertain. Recent studies have examined LLM use in various health care applications, but their effectiveness for providing nutrition information remains understudied. Although retrieval-augmented generation (RAG) frameworks have been shown to enhance LLM consistency and accuracy, their use in delivering nutrition information for CVD prevention requires further evaluation. Objective: To evaluate the effectiveness of off-the-shelf and RAG-enhanced LLMs in delivering guideline-adherent nutrition information for CVD prevention, we assessed 3 off-the-shelf models (ChatGPT-4o, Perplexity, and Llama 3-70B) and a Llama 3-70B+RAG model. Methods: We curated 30 nutrition questions that comprehensively addressed CVD prevention. These were approved by a registered dietitian providing preventive cardiology services at an academic medical center and were posed 3 times to each model. We developed a 15,074-word knowledge bank incorporating the American Heart Association’s 2021 dietary guidelines and related website content to enhance Meta’s Llama 3-70B model using RAG. The model received this and a few-shot prompt as context, included citations in a Context Source section, and used vector similarity to align responses with guideline content, with the temperature parameter set to 0.5 to enhance consistency. Model responses were evaluated by 3 expert reviewers against benchmark CV guidelines for appropriateness, reliability, readability, harm, and guideline adherence. Mean scores were compared using ANOVA, with statistical significance set at P<.05. interrater agreement was measured using the cohen coefficient and readability estimated flesch-kincaid score. results: llama model scored higher than perplexity gpt-4o models on reliability appropriateness guideline adherence showed no harm.>70%; P<.001 indicated high reviewer agreement. conclusions: the llama model outperformed off-the-shelf models across all measures with no evidence of harm although responses were less readable due to technical language. scored lower on and produced some harmful responses. these findings highlight limitations demonstrate that rag system integration can enhance llm performance in delivering evidence-based dietary information.>

The Role of Data in Public Health and Health Innovation: Perspectives on Social Determinants of Health, Community-Based Data Approaches, and AI

Public health is undergoing profound transformation driven by data from the global health sector and related fields. To address systemic health disparities, scholars and practitioners are increasingly applying a data equity lens, an approach that has become even more urgent as the United States faces the erosion of public health data infrastructure. This paper summarizes insights from an April 2024 convening by the Yale School of Public Health—The Role of Data in Public Health Equity and Innovation—with intersectoral stakeholders from academia, government (local, state, and federal), healthcare, and private industry. The convening included keynote presentations and roundtables regarding the depiction of social determinants of health (SDOH) in data; effects of artificial intelligence (AI) on health data equity; and community-based models for data, providing a framework for cross-cutting discussions. Through a narrative synthesis, themes were identified and synthesized from systematically gathered information from presentations and roundtables. This process led to a set of actionable, cross-cutting recommendations to guide inclusive and impactful data practices for policymakers, public health professionals, and health innovators across diverse contexts: (1) Enable big data and interoperability connecting SDOH and health outcomes; (2) Include diverse, non-technical voices in AI and health discussions; (3) Fund research on data equity and AI in health sciences; (4) Modernize Health Insurance Portability and Accountability Act (HIPAA) with new guidelines for AI and big data; and (5) Research and conceptual frameworks are needed to elucidate interconnections between data equity and health equity.

Generative artificial intelligence in medicine

Nature Medicine, Published online: 06 October 2025; doi:10.1038/s41591-025-03983-2

This Review summarizes recent technical advancements in generative AI, outlines how new models might improve healthcare and discusses validation approaches—using lessons from recent successes and failures in the field.

Exploring Attitudes and Obstacles Around Digital Public Health Tools: Insights From a Statewide Cross-Sectional Survey on Washington’s Vaccine Verification System

Background: Development and use of digital public health tools surged during the COVID-19 pandemic. Among these tools, vaccine verification systems emerged as alternatives to paper vaccine records, aiming to help limit the spread of disease. In November 2021, the Washington State Department of Health launched “WA Verify,” a QR code–based vaccine verification system built on the SMART Health Card framework, providing residents with a convenient way to store and share proof of vaccination digitally. However, WA Verify was developed and deployed before assessments and public input regarding potential adoption challenges—such as concerns about privacy, surveillance, data sharing, trust in the technology, and the managing organizations—could be completed. Objective: This analysis used statewide survey data from Washington to identify and characterize barriers and facilitators to the adoption of WA Verify, and to understand how factors such as data privacy, security, attitudes toward public health policies and communication, and technological proficiency may influence acceptance and uptake of digital public health tools. Methods: A cross-sectional statewide survey was distributed between September 2022 and January 2023 to a random sample of 5000 Washington households. Respondents were categorized into 3 groups based on their responses indicating WA Verify “users,” “potential users,” or “unlikely users.” Comparisons were made between groups regarding experiences with and opinions on COVID-19 vaccine and test verification, public health policies, communication, digital tools, technological proficiency, sociodemographic characteristics, and health history. Poststratification weights were applied to reduce nonresponse bias. Results: Of the 1401 respondents, 359 (25.6% unweighted, 25.8% weighted) were users, 662 (47.3% unweighted, 49.8% weighted) were potential users, and 380 (27.1% unweighted, 24.4% weighted) were unlikely users. All percentages reported are based on weighted data. Compared with users and potential users, unlikely users were more likely to oppose policies requiring proof of COVID-19 vaccination or negative test results (users: 6.0%, potential users: 13.6%, unlikely users: 65.9%). Unlikely users were more likely to cite concerns about personal health data security and phone hacking or tracking, though these concerns were also notable among potential users and users. Users and potential users were more likely to perceive a digital vaccine verification system as convenient (users: 96.5%, potential users: 92.3%, unlikely users: 38.1%) and indicated openness to receiving relevant information from a range of sources. Unlikely users were more likely to report not owning a smartphone and demonstrated lower technological proficiency (users: 12.3%, potential users: 15.9%, unlikely users: 32.3%), indicating a technological divide between groups. Conclusions: While nearly three-quarters of respondents had either already adopted or were willing to adopt a tool like WA Verify, concerns about data security, lower technological proficiency, and distrust of public health characterized those least likely to adopt such tools. Identifying barriers to adoption among “unlikely users” is essential for developing effective communication strategies—such as targeted marketing and community engagement—to improve adoption and ensure equitable access to public health technologies.

Comparative Evaluation of a Medical Large Language Model in Answering Real-World Radiation Oncology Questions: Multicenter Observational Study

Background: Large language models (LLMs) hold promise for supporting clinical tasks, particularly in data-driven and technical disciplines such as radiation oncology. While prior evaluation studies have focused on examination-style settings for evaluating LLMs, their performance in real-life clinical scenarios remains unclear. In the future, LLMs might be used as general AI assistants to answer questions arising in clinical practice. It is unclear how well a modern LLM, locally executed within the infrastructure of a hospital, would answer such questions compared with clinical experts. Objective: This study aimed to assess the performance of a locally deployed, state-of-the-art medical LLM in answering real-world clinical questions in radiation oncology compared with clinical experts. The aim was to evaluate the overall quality of answers, as well as the potential harmfulness of the answers if used for clinical decision-making. Methods: Physicians from 10 departments of European hospitals collected questions arising in the clinical practice of radiation oncology. Fifty of these questions were answered by 3 senior radiation oncology experts with at least 10 years of work experience, as well as the LLM Llama3-OpenBioLLM-70B (Ankit Pal and Malaikannan Sankarasubbu). In a blinded review, physicians rated the overall answer quality on a 5-point Likert scale (quality), assessed whether an answer might be potentially harmful if used for clinical decision-making (harmfulness), and determined if responses were from an expert or the LLM (recognizability). Comparisons between clinical experts and LLMs were then made for quality, harmfulness, and recognizability. Results: There were no significant differences between the quality of the answers between LLM and clinical experts (mean scores of 3.38 vs 3.63; median 4.00, IQR 3.00-4.00 vs median 3.67, IQR 3.33-4.00; P=.26; Wilcoxon signed rank test). The answers were deemed potentially harmful in 13% of cases for the clinical experts compared with 16% of cases for the LLM (P=.63; Fisher exact test). Physicians correctly identified whether an answer was given by a clinical expert or an LLM in 78% and 72% of cases, respectively. Conclusions: A state-of-the-art medical LLM can answer real-life questions from the clinical practice of radiation oncology similarly well as clinical experts regarding overall quality and potential harmfulness. Such LLMs can already be deployed within the local hospital environment at an affordable cost. While LLMs may not yet be ready for clinical implementation as general AI assistants, the technology continues to improve at a rapid pace. Evaluation studies based on real-life situations are important to better understand the weaknesses and limitations of LLMs in clinical practice. Such studies are also crucial to define when the technology is ready for clinical implementation. Furthermore, education for health care professionals on generative AI is needed to ensure responsible clinical implementation of this transforming technology.

Large Language Models’ Clinical Decision-Making on When to Perform a Kidney Biopsy: Comparative Study

Background: Artificial intelligence (AI) and Large Language models (LLMs) are increasing in sophistication and are being integrated into many disciplines. The potential for LLMs to augment clinical decisions is an evolving area of research. Objective: This study compared the responses of over 1000 kidney specialist physicians (nephrologists) to outputs of commonly used LLMs using a questionnaire determining when a kidney biopsy should be performed. Methods: This research group completed a large online questionnaire for nephrologists to determine when a kidney biopsy should be performed. The questionnaire was co-designed with patient participation, refined through multiple iterations, then piloted locally before international dissemination. It was the largest international study in the field and demonstrated variation between human clinicians in biopsy propensity relating to human factors such as sex and age, as well as systemic factors such as country, job seniority and technical proficiency. The same questions were put to both human doctors and LLMs in an identical order in a single session. Eight commonly used LLMs were interrogated: Chat GPT 3.5, Mistral Hugging Face, Perplexity, Microsoft Co-pilot, Llama 2, GPT 4.0, MedLM and Claude 3. The most common response given by clinicians (human mode) to each question was taken as the baseline for comparison. Questionnaire responses to the indications and contraindications for biopsy generated a score (0-44) reflecting biopsy propensity, in which a higher score was used as a surrogate marker for an increased tolerance of potential associated risks. Results: The ability of LLMs to reproduce human expert consensus varied widely with some models demonstrating a balanced approach to risk in a similar manner to humans, whilst other models reported outputs at either end of the spectrum for risk tolerance. In terms of agreement with the human mode, Chat GPT 3.5 and GPT 4.0 (Open AI) had the highest levels of alignment, with the human mode selected in 6/11 questions. The total biopsy propensity score generated from the human mode was 23/44. Both Open AI models produced similar propensity scores between 22 and 24, however Llama 2 and MS Co-pilot also reported scores within this range, but with poorer response alignment to the human mode at only 2/11 questions. The most risk averse model in this study was MedLM with a propensity score of 11 and the least risk averse model was Claude 3 with a score of 34. Conclusions: LLM outputs demonstrated a modest ability to replicate human clinical decision making in this study, however the performance varied widely between LLM models. Questions with more uniform human responses produced LLM outputs with greater alignment, whereas in questions with low levels of human consensus there was poor output alignment. This may limit the practical use of LLMs in real world clinical practice.
❌