❌

Normal view

Stakeholder Criteria for Trust in Artificial Intelligence–Based Computer Perception Tools in Health Care: Qualitative Interview Study

Background: Computer perception (CP) technologies hold significant promise for advancing precision mental health care systems, given their ability to leverage algorithmic analysis of continuous, passive sensing data from wearables and smartphones (eg, behavioral activity, geolocation, vocal features, and ambient environmental data) to infer clinically meaningful behavioral and physiological states. However, successful implementation critically depends on cultivating well-founded stakeholder trust. Objective: This study aims to investigate, across adolescents, caregivers, clinicians, and developers, the contingencies under which CP technologies are deemed trustworthy in health care. Methods: We conducted 80 semistructured interviews with a purposive sample of adolescents (n=20) diagnosed with autism, Tourette syndrome, anxiety, obsessive-compulsive disorder, or attention-deficit/hyperactivity disorder and their caregivers (n=20); practicing clinicians across psychiatry, psychology, and pediatrics (n=20); and CP system developers (n=20). Interview transcripts were coded by 2 independent coders and analyzed using multistage, inductive thematic content analysis to identify prominent themes. Results: Across stakeholder groups, 5 core criteria emerged as prerequisites for trust in CP outputs: (1) epistemic alignment—consistency between system outputs, personal experience, and existing diagnostic frameworks; (2) demonstrable rigor—training on representative data and validation in real-world contexts; (3) explainability—transparent communication of input variables, thresholds, and decision logic; (4) sensitivity to complexity—the capacity to accommodate heterogeneity and comorbidity in symptom expression; and (5) a nonsubstitutive role—technologies must augment, rather than supplant, clinical judgment. A novel and cautionary finding was that epistemic alignment—whether outputs affirmed participants’ preexisting beliefs, diagnostic expectations, or internal states—was a dominant factor in determining whether the tool was perceived as trustworthy. Participants also expressed relational trust, placing confidence in CP systems based on endorsements from respected peers, academic institutions, or regulatory agencies. However, both trust strategies raise significant concerns: confirmation bias may lead users to overvalue outputs that align with their assumptions, while surrogate trust may be misapplied in the absence of robust performance validation. Conclusions: This study advances empirical understanding of how trust is formed and calibrated around artificial intelligence–based CP technologies. While trust is commonly framed as a function of technical performance, our findings show that it is deeply shaped by cognitive heuristics, social relationships, and alignment with entrenched epistemologies. These dynamics can facilitate intuitive verification but may also constrain the transformative potential of CP systems by reinforcing existing beliefs. To address this, we recommend a dual strategy: (1) embedding CP tools within institutional frameworks that uphold rigorous validation, ethical oversight, and transparent design; and (2) providing clinicians with training and interface designs that support critical appraisal and minimize susceptibility to cognitive bias. Recalibrating trust to reflect actual system capacities—rather than familiarity or endorsement—is essential for ethically sound and clinically meaningful integration of CP technologies.

Trump’s AI executive order promises ‘one rulebook’ — startups may get legal limbo instead

13 December 2025 at 01:07
Trump signed an AI executive order targeting state laws and promising one national rulebook. Critics warn it could trigger court battles and prolong uncertainty for startups while Congress debates federal rules.

Exploring Health Misinformation Detection with Multi-Agent Debate

arXiv:2512.09935v1 Announce Type: new Abstract: Fact-checking health-related claims has become increasingly critical as misinformation proliferates online. Effective verification requires both the retrieval of high-quality evidence and rigorous reasoning processes. In this paper, we propose a two-stage framework for health misinformation detection: Agreement Score Prediction followed by Multi-Agent Debate. In the first stage, we employ large language models (LLMs) to independently evaluate retrieved articles and compute an aggregated agreement score that reflects the overall evidence stance. When this score indicates insufficient consensus-falling below a predefined threshold-the system proceeds to a second stage. Multiple agents engage in structured debate to synthesize conflicting evidence and generate well-reasoned verdicts with explicit justifications. Experimental results demonstrate that our two-stage approach achieves superior performance compared to baseline methods, highlighting the value of combining automated scoring with collaborative reasoning for complex verification tasks.

Mind the Gap! Pathways Towards Unifying AI Safety and Ethics Research

arXiv:2512.10058v1 Announce Type: new Abstract: While much research in artificial intelligence (AI) has focused on scaling capabilities, the accelerating pace of development makes countervailing work on producing harmless, "aligned" systems increasingly urgent. Yet research on alignment has diverged along two largely parallel tracks: safety--centered on scaled intelligence, deceptive or scheming behaviors, and existential risk--and ethics--focused on present harms, the reproduction of social bias, and flaws in production pipelines. Although both communities warn of insufficient investment in alignment, they disagree on what alignment means or ought to mean. As a result, their efforts have evolved in relative isolation, shaped by distinct methodologies, institutional homes, and disciplinary genealogies. We present a large-scale, quantitative study showing the structural split between AI safety and AI ethics. Using a bibliometric and co-authorship network analysis of 6,442 papers from twelve major ML and NLP conferences (2020-2025), we find that over 80% of collaborations occur within either the safety or ethics communities, and cross-field connectivity is highly concentrated: roughly 5% of papers account for more than 85% of bridging links. Removing a small number of these brokers sharply increases segregation, indicating that cross-disciplinary exchange depends on a handful of actors rather than broad, distributed collaboration. These results show that the safety-ethics divide is not only conceptual but institutional, with implications for research agendas, policy, and venues. We argue that integrating technical safety work with normative ethics--via shared benchmarks, cross-institutional venues, and mixed-method methodologies--is essential for building AI systems that are both robust and just.

Are ultrasensitive ctDNA assays ready for clinical use in early-stage NSCLC?

11 December 2025 at 08:00
Disease recurrence in early-stage non-small cell lung cancer (NSCLC) remains a persistent clinical challenge, underscoring the need for better prognostic biomarkers. In this preview, we highlight the clinical implications of ultrasensitive ctDNA monitoring in lung cancer risk modeling reported by Black et al. in this issue of Cell.

Macrophage-targeted immunocytokine leverages myeloid, T, and NK cell synergy for cancer immunotherapy

MiTEs are myeloid-targeted immunocytokine prodrugs that block TREM2+ tumor-associated macrophages while activating cytotoxic lymphocytes via TME-specific IL-2 activity, eliciting strong anti-tumor efficacy in preclinical models with minimal systemic toxicity.

Pancreatic Cancer Organoids: Modeling Disease and Guiding Therapy

Cancers (Basel). 2025 Nov 30;17(23):3850. doi: 10.3390/cancers17233850.

ABSTRACT

Pancreatic ductal adenocarcinoma (PDAC) is one of the most lethal malignancies. An unmet need exists for reliable biomarkers and in vitro models capable of predicting patient drug response to advance personalized medicine. Traditional models fail to represent the tumor's complexity and the role of the stromal environment in chemoresistance. Patient-derived organoids (PDOs) overcome these limitations, enabling multi-omics profiling and reliable drug testing for functional precision medicine. This review provides a comprehensive overview of PDAC PDO research, emphasizing the following major areas: (i) the genetic and phenotypic fidelity of PDOs, (ii) their predictive value for drug response and chemoresistance, (iii) the integration of the extracellular matrix and tumor microenvironment (TME) components, and (iv) emerging technologies. Studies confirm that PDOs faithfully represent the primary tumor's specific genetic features and retain intratumoral heterogeneity. PDO-based platforms have demonstrated a strong correlation between in vitro drug sensitivity and in vivo efficacy in xenograft models, validating their utility for identifying drug candidates, repurposing existing drugs, and determining effective combinations. Efforts are ongoing to integrate crucial TME components, like cancer-associated fibroblasts, using innovative co-culture platforms such as fused PDOs and InterOMaX, to better model desmoplasia and chemoresistance mechanisms. Furthermore, PDO technology is converging with microphysiological systems and artificial intelligence tools to facilitate high-throughput drug screening and dynamic, real-time monitoring of therapeutic effects. The integration of PDOs into biobanks and advanced screening platforms holds the potential to accelerate drug discovery and improve therapeutic outcomes for PDAC patients, if challenges related to protocol standardization and regulatory acceptance are addressed.

PMID:41375051 | PMC:PMC12690986 | DOI:10.3390/cancers17233850

Integrative network pharmacology, transcriptomics, and microbiomics elucidate the therapeutic mechanism of <em>Polygala tenuifolia</em> Willd water extract in chronic obstructive pulmonary disease

11 December 2025 at 19:00

Front Microbiol. 2025 Nov 25;16:1703853. doi: 10.3389/fmicb.2025.1703853. eCollection 2025.

ABSTRACT

BACKGROUND: Polygala tenuifolia Willd (PT) is a plant with both medicinal and edible values. Traditionally, it has been used for sedation, enhancing cognition, resolving phlegm, and relieving cough. However, its protective effects and mechanisms against chronic obstructive pulmonary disease (COPD) remain unclear.

AIM OF THE STUDY: This study aims to observe the protective effects of the water extract of Polygala tenuifolia Willd (WEPT) on COPD, and to preliminarily elucidate its potential therapeutic mechanisms by integrating network pharmacology, molecular docking, multi-omics analysis, and molecular experiments.

METHODS AND MATERIALS: HPLC quantified WEPT constituents. COPD mice models established via chronic smoke exposure underwent WEPT treatment, and the therapeutic effect was evaluated by lung function test, histopathology and cytokine profiling. Integrated multi-omics analyses (network pharmacology, transcriptomics, microbiomics) identified bioactive compounds, therapeutic targets, pathway regulations, and microbiota dynamics. Molecular docking validated compound-target interactions, while immunohistochemical/fluorescence assays confirmed key protein expression in lung tissues.

RESULTS: WEPT administration effectively reduced inflammatory cytokine levels in COPD mice, improved lung function, and alleviated histopathological damage like alveolar structural injury and airway inflammation. Network pharmacology and transcriptomic analyses identified Norhyoscyamine and Onjixanthone I as key active components, targeting PIK3CA and AKT1 via PI3K-AKT pathway regulation. Microbiome analysis showed WEPT restored gut microbiota balance. Molecular docking confirmed strong binding of bioactive compounds to core targets, while immunostaining assays demonstrated WEPT suppressed p-PI3K and p-AKT protein expression.

CONCLUSION: WEPT may exert its intervention effects on COPD through a multi-target and multi-level comprehensive regulatory mechanism.

PMID:41377050 | PMC:PMC12685879 | DOI:10.3389/fmicb.2025.1703853

AI-driven virtual cell models in preclinical research: technical pathways, validation mechanisms, and clinical translation potential

npj Digital Medicine, Published online: 11 December 2025; doi:10.1038/s41746-025-02198-6

AI-driven virtual cell models in preclinical research: technical pathways, validation mechanisms, and clinical translation potential

Toward an AI Reasoning-Enabled System for Patient-Clinical Trial Matching

arXiv:2512.08026v1 Announce Type: new Abstract: Screening patients for clinical trial eligibility remains a manual, time-consuming, and resource-intensive process. We present a secure, scalable proof-of-concept system for Artificial Intelligence (AI)-augmented patient-trial matching that addresses key implementation challenges: integrating heterogeneous electronic health record (EHR) data, facilitating expert review, and maintaining rigorous security standards. Leveraging open-source, reasoning-enabled large language models (LLMs), the system moves beyond binary classification to generate structured eligibility assessments with interpretable reasoning chains that support human-in-the-loop review. This decision support tool represents eligibility as a dynamic state rather than a fixed determination, identifying matches when available and offering actionable recommendations that could render a patient eligible in the future. The system aims to reduce coordinator burden, intelligently broaden the set of trials considered for each patient and guarantee comprehensive auditability of all AI-generated outputs.

Beyond Traditional Diagnostics: Transforming Patient-Side Information into Predictive Insights with Knowledge Graphs and Prototypes

arXiv:2512.08261v1 Announce Type: new Abstract: Predicting diseases solely from patient-side information, such as demographics and self-reported symptoms, has attracted significant research attention due to its potential to enhance patient awareness, facilitate early healthcare engagement, and improve healthcare system efficiency. However, existing approaches encounter critical challenges, including imbalanced disease distributions and a lack of interpretability, resulting in biased or unreliable predictions. To address these issues, we propose the Knowledge graph-enhanced, Prototype-aware, and Interpretable (KPI) framework. KPI systematically integrates structured and trusted medical knowledge into a unified disease knowledge graph, constructs clinically meaningful disease prototypes, and employs contrastive learning to enhance predictive accuracy, which is particularly important for long-tailed diseases. Additionally, KPI utilizes large language models (LLMs) to generate patient-specific, medically relevant explanations, thereby improving interpretability and reliability. Extensive experiments on real-world datasets demonstrate that KPI outperforms state-of-the-art methods in predictive accuracy and provides clinically valid explanations that closely align with patient narratives, highlighting its practical value for patient-centered healthcare delivery.

Principles2Plan: LLM-Guided System for Operationalising Ethical Principles into Plans

arXiv:2512.08536v1 Announce Type: new Abstract: Ethical awareness is critical for robots operating in human environments, yet existing automated planning tools provide little support. Manually specifying ethical rules is labour-intensive and highly context-specific. We present Principles2Plan, an interactive research prototype demonstrating how a human and a Large Language Model (LLM) can collaborate to produce context-sensitive ethical rules and guide automated planning. A domain expert provides the planning domain, problem details, and relevant high-level principles such as beneficence and privacy. The system generates operationalisable ethical rules consistent with these principles, which the user can review, prioritise, and supply to a planner to produce ethically-informed plans. To our knowledge, no prior system supports users in generating principle-grounded rules for classical planning contexts. Principles2Plan showcases the potential of human-LLM collaboration for making ethical automated planning more practical and feasible.

Multi-Agent Intelligence for Multidisciplinary Decision-Making in Gastrointestinal Oncology

arXiv:2512.08674v1 Announce Type: new Abstract: Multimodal clinical reasoning in the field of gastrointestinal (GI) oncology necessitates the integrated interpretation of endoscopic imagery, radiological data, and biochemical markers. Despite the evident potential exhibited by Multimodal Large Language Models (MLLMs), they frequently encounter challenges such as context dilution and hallucination when confronted with intricate, heterogeneous medical histories. In order to address these limitations, a hierarchical Multi-Agent Framework is proposed, which emulates the collaborative workflow of a human Multidisciplinary Team (MDT). The system attained a composite expert evaluation score of 4.60/5.00, thereby demonstrating a substantial improvement over the monolithic baseline. It is noteworthy that the agent-based architecture yielded the most substantial enhancements in reasoning logic and medical accuracy. The findings indicate that mimetic, agent-based collaboration provides a scalable, interpretable, and clinically robust paradigm for automated decision support in oncology.

Towards Foundation Models with Native Multi-Agent Intelligence

arXiv:2512.08743v1 Announce Type: new Abstract: Foundation models (FMs) are increasingly assuming the role of the "brain" of AI agents. While recent efforts have begun to equip FMs with native single-agent abilities -- such as GUI interaction or integrated tool use -- we argue that the next frontier is endowing FMs with native multi-agent intelligence. We identify four core capabilities of FMs in multi-agent contexts: understanding, planning, efficient communication, and adaptation. Contrary to assumptions about the spontaneous emergence of such abilities, we provide extensive empirical evidence across 41 large language models showing that strong single-agent performance alone does not automatically yield robust multi-agent intelligence. To address this gap, we outline key research directions -- spanning dataset construction, evaluation, training paradigms, and safety considerations -- for building FMs with native multi-agent intelligence.

Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture

arXiv:2512.08130v1 Announce Type: cross Abstract: Both model developers and policymakers seek to quantify and mitigate the risk of rapidly-evolving frontier artificial intelligence (AI) models, especially large language models (LLMs), to facilitate bioterrorism or access to biological weapons. An important element of such efforts is the development of model benchmarks that can assess the biosecurity risk posed by a particular model. This paper describes the first component of a novel Biothreat Benchmark Generation (BBG) Framework. The BBG approach is designed to help model developers and evaluators reliably measure and assess the biosecurity risk uplift and general harm potential of existing and future AI models, while accounting for key aspects of the threat itself that are often overlooked in other benchmarking efforts, including different actor capability levels, and operational (in addition to purely technical) risk factors. As a pilot, the BBG is first being developed to address bacterial biological threats only. The BBG is built upon a hierarchical structure of biothreat categories, elements and tasks, which then serves as the basis for the development of task-aligned queries. This paper outlines the development of this biothreat task-query architecture, which we have named the Bacterial Biothreat Schema, while future papers will describe follow-on efforts to turn queries into model prompts, as well as how the resulting benchmarks can be implemented for model evaluation. Overall, the BBG Framework, including the Bacterial Biothreat Schema, seeks to offer a robust, re-usable structure for evaluating bacterial biological risks arising from LLMs across multiple levels of aggregation, which captures the full scope of technical and operational requirements for biological adversaries, and which accounts for a wide spectrum of biological adversary capabilities.

A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties

arXiv:2512.08185v1 Announce Type: cross Abstract: Medical Large Language Models (LLMs) are increasingly deployed for clinical decision support across diverse specialties, yet systematic evaluation of their robustness to adversarial misuse and privacy leakage remains inaccessible to most researchers. Existing security benchmarks require GPU clusters, commercial API access, or protected health data -- barriers that limit community participation in this critical research area. We propose a practical, fully reproducible framework for evaluating medical AI security under realistic resource constraints. Our framework design covers multiple medical specialties stratified by clinical risk -- from high-risk domains such as emergency medicine and psychiatry to general practice -- addressing jailbreaking attacks (role-playing, authority impersonation, multi-turn manipulation) and privacy extraction attacks. All evaluation utilizes synthetic patient records requiring no IRB approval. The framework is designed to run entirely on consumer CPU hardware using freely available models, eliminating cost barriers. We present the framework specification including threat models, data generation methodology, evaluation protocols, and scoring rubrics. This proposal establishes a foundation for comparative security assessment of medical-specialist models and defense mechanisms, advancing the broader goal of ensuring safe and trustworthy medical AI systems.

ClinicalTrialsHub: Bridging Registries and Literature for Comprehensive Clinical Trial Access

arXiv:2512.08193v1 Announce Type: cross Abstract: We present ClinicalTrialsHub, an interactive search-focused platform that consolidates all data from ClinicalTrials.gov and augments it by automatically extracting and structuring trial-relevant information from PubMed research articles. Our system effectively increases access to structured clinical trial data by 83.8% compared to relying on ClinicalTrials.gov alone, with potential to make access easier for patients, clinicians, researchers, and policymakers, advancing evidence-based medicine. ClinicalTrialsHub uses large language models such as GPT-5.1 and Gemini-3-Pro to enhance accessibility. The platform automatically parses full-text research articles to extract structured trial information, translates user queries into structured database searches, and provides an attributed question-answering system that generates evidence-grounded answers linked to specific source sentences. We demonstrate its utility through a user study involving clinicians, clinical researchers, and PhD students of pharmaceutical sciences and nursing, and a systematic automatic evaluation of its information extraction and question answering capabilities.

Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models III: Implementing the Bacterial Biothreat Benchmark (B3) Dataset

arXiv:2512.08459v1 Announce Type: cross Abstract: The potential for rapidly-evolving frontier artificial intelligence (AI) models, especially large language models (LLMs), to facilitate bioterrorism or access to biological weapons has generated significant policy, academic, and public concern. Both model developers and policymakers seek to quantify and mitigate any risk, with an important element of such efforts being the development of model benchmarks that can assess the biosecurity risk posed by a particular model. This paper discusses the pilot implementation of the Bacterial Biothreat Benchmark (B3) dataset. It is the third in a series of three papers describing an overall Biothreat Benchmark Generation (BBG) framework, with previous papers detailing the development of the B3 dataset. The pilot involved running the benchmarks through a sample frontier AI model, followed by human evaluation of model responses, and an applied risk analysis of the results along several dimensions. Overall, the pilot demonstrated that the B3 dataset offers a viable, nuanced method for rapidly assessing the biosecurity risk posed by a LLM, identifying the key sources of that risk and providing guidance for priority areas of mitigation priority.
❌