❌

Normal view

Exploring Health Misinformation Detection with Multi-Agent Debate

arXiv:2512.09935v1 Announce Type: new Abstract: Fact-checking health-related claims has become increasingly critical as misinformation proliferates online. Effective verification requires both the retrieval of high-quality evidence and rigorous reasoning processes. In this paper, we propose a two-stage framework for health misinformation detection: Agreement Score Prediction followed by Multi-Agent Debate. In the first stage, we employ large language models (LLMs) to independently evaluate retrieved articles and compute an aggregated agreement score that reflects the overall evidence stance. When this score indicates insufficient consensus-falling below a predefined threshold-the system proceeds to a second stage. Multiple agents engage in structured debate to synthesize conflicting evidence and generate well-reasoned verdicts with explicit justifications. Experimental results demonstrate that our two-stage approach achieves superior performance compared to baseline methods, highlighting the value of combining automated scoring with collaborative reasoning for complex verification tasks.

A Field Guide to Deploying AI Agents in Clinical Practice

arXiv:2509.26153v3 Announce Type: replace Abstract: Large language models (LLMs) integrated into agent-driven workflows hold immense promise for healthcare, yet a significant gap exists between their potential and practical implementation within clinical settings. To address this, we present a practitioner-oriented field manual for deploying generative agents that use electronic health record (EHR) data. This guide is informed by our experience deploying the "irAE-Agent", an automated system to detect immune-related adverse events from clinical notes at Mass General Brigham, and by structured interviews with 21 clinicians, engineers, and informatics leaders involved in the project. Our analysis reveals a critical misalignment in clinical AI development: less than 20% of our effort was dedicated to prompt engineering and model development, while over 80% was consumed by the sociotechnical work of implementation. We distill this effort into five "heavy lifts": data integration, model validation, ensuring economic value, managing system drift, and governance. By providing actionable solutions for each of these challenges, this field manual shifts the focus from algorithmic development to the essential infrastructure and implementation work required to bridge the "valley of death" and successfully translate generative AI from pilot projects into routine clinical care.

From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems

arXiv:2503.01424v4 Announce Type: replace Abstract: Research is a fundamental process driving the advancement of human civilization, yet it demands substantial time and effort from researchers. In recent years, the rapid development of artificial intelligence (AI) technologies has inspired researchers to explore how AI can accelerate and enhance research. To monitor relevant advancements, this paper presents a systematic review of the progress in this domain. Specifically, we organize the relevant studies into three main categories: hypothesis formulation, hypothesis validation, and manuscript publication. Hypothesis formulation involves knowledge synthesis and hypothesis generation. Hypothesis validation includes the verification of scientific claims, theorem proving, and experiment validation. Manuscript publication encompasses manuscript writing and the peer review process. Furthermore, we identify and discuss the current challenges faced in these areas, as well as potential future directions for research. Finally, we also offer a comprehensive overview of existing benchmarks and tools across various domains that support the integration of AI into the research process. We hope this paper serves as an introduction for beginners and fosters future research. Resources have been made publicly available at https://github.com/zkzhou126/AI-for-Research.

MARIA: A Framework for Marginal Risk Assessment without Ground Truth in AI Systems

arXiv:2510.27163v1 Announce Type: cross Abstract: Before deploying an AI system to replace an existing process, it must be compared with the incumbent to ensure improvement without added risk. Traditional evaluation relies on ground truth for both systems, but this is often unavailable due to delayed or unknowable outcomes, high costs, or incomplete data, especially for long-standing systems deemed safe by convention. The more practical solution is not to compute absolute risk but the difference between systems. We therefore propose a marginal risk assessment framework, that avoids dependence on ground truth or absolute risk. It emphasizes three kinds of relative evaluation methodology, including predictability, capability and interaction dominance. By shifting focus from absolute to relative evaluation, our approach equips software teams with actionable guidance: identifying where AI enhances outcomes, where it introduces new risks, and how to adopt such systems responsibly.

From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learning

arXiv:2510.24528v1 Announce Type: new Abstract: The capability of in-context learning (ICL) enables large language models (LLMs) to perform novel tasks without parameter updates by conditioning on a few input-output examples. However, collecting high-quality examples for new or challenging tasks can be costly and labor-intensive. In this work, we propose a cost-efficient two-stage pipeline that reduces reliance on LLMs for data labeling. Our approach first leverages readily available cross-task examples to prompt an LLM and pseudo-label a small set of target task instances. We then introduce a graph-based label propagation method that spreads label information to the remaining target examples without additional LLM queries. The resulting fully pseudo-labeled dataset is used to construct in-task demonstrations for ICL. This pipeline combines the flexibility of cross-task supervision with the scalability of LLM-free propagation. Experiments across five tasks demonstrate that our method achieves strong performance while lowering labeling costs.

When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior

npj Digital Medicine, Published online: 17 October 2025; doi:10.1038/s41746-025-02008-z

When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior

Fine-Tuning Methods for Large Language Models in Clinical Medicine by Supervised Fine-Tuning and Direct Preference Optimization: Comparative Evaluation

Background: Large language model (LLM) fine tuning is the process of adjusting out-of-the-box model weights using a dataset of interest. Fine tuning can be a powerful technique to improve model performance in fields like medicine, where data access is restricted and LLMs may have poor out-of-the-box performance. Objective: In this study we investigated the benefits of fine tuning with supervised fine tuning (SFT) and direct preference optimization (DPO) across a range of LLM applications for medicine Methods: We use Llama3 7B and Mistral 7B v2 to compare the performance of SFT and DPO across four datasets for common natural language tasks in medicine. The tasks evaluated were simple classification, clinical reasoning, summarization, and clinical triage. Results: Clinical Reasoning accuracy increased 8% and 7% with DPO over SFT for Llama3 (p value 0.003) and Mistral2 (p value 0.004) respectively. Summarization quality, graded on a five point Likert scale, increased 0.13 and 0.10 for Llama3 and Mistral2 (p values

A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies

Nat Comput Sci. 2025 Feb 7. doi: 10.1038/s43588-024-00764-8. Online ahead of print.

ABSTRACT

Large-scale whole-genome sequencing (WGS) studies have improved our understanding of the contributions of coding and noncoding rare variants to complex human traits. Leveraging association effect sizes across multiple traits in WGS rare variant association analysis can improve statistical power over single-trait analysis, and also detect pleiotropic genes and regions. Existing multi-trait methods have limited ability to perform rare variant analysis of large-scale WGS data. We propose MultiSTAAR, a statistical framework and computationally scalable analytical pipeline for functionally informed multi-trait rare variant analysis in large-scale WGS studies. MultiSTAAR accounts for relatedness, population structure and correlation among phenotypes by jointly analyzing multiple traits, and further empowers rare variant association analysis by incorporating multiple functional annotations. We applied MultiSTAAR to jointly analyze three lipid traits in 61,838 multi-ethnic samples from the Trans-Omics for Precision Medicine (TOPMed) Program. We discovered and replicated new associations with lipid traits missed by single-trait analysis.

PMID:39920506 | DOI:10.1038/s43588-024-00764-8

Integrating multiomics analysis and machine learning to refine the molecular subtyping and prognostic analysis of stomach adenocarcinoma

Sci Rep. 2025 Jan 30;15(1):3843. doi: 10.1038/s41598-025-87444-3.

ABSTRACT

Stomach adenocarcinoma (STAD) is a common malignancy with high heterogeneity and a lack of highly precise treatment options. We downloaded the multiomics data of STAD patients in The Cancer Genome Atlas (TCGA)-STAD cohort, which included mRNA, microRNA, long non-coding RNA, somatic mutation, and DNA methylation data, from the sxdyc website. We synthesized the multiomics data of patients with STAD using 10 clustering methods, construct a consensus machine learning-driven signature (CMLS)-related prognostic models by combining 10 machine learning methods, and evaluated the prognosis models using the C-index. The prognostic relationship between CMLS and STAD was assessed using Kaplan-Meier curves, and the independent prognostic value of CMLS was determined by univariate and multivariate regression analyses. we also evaluated the immune characteristics, immunotherapy response, and drug sensitivity of different CMLS groups. The results of the multiomics analysis classified STAD into three subtypes, with CS1 resulting in the best survival outcome. In total, 10 hub genes (CES3, AHCYL2, APOD, EFEMP1, CYP1B1, ASPN, CPE, CLIP3, MAP1B, and DKK1) were screened and constructed the CMLS was significantly correlated with prognosis in patients with STAD and was an independent prognostic factor for patients with STAD. Using the CMLS risk score, all patients were divided into a high CMLS group and a low CMLS group. Patients in the low-CMLS group had better survival, more enriched immune cells, and higher tumor mutation load scores, suggesting better immunotherapy responsiveness and a possible "hot tumor" phenotype. Patients in the high-CMLS group had a significantly poorer prognosis and were less sensitive to immunotherapy but were likely to benefit more from chemotherapy and targeted therapy. In this study, 10 clustering methods and 10 machine learning methods were combined to analyze the multiomics of STAD, classify STAD into three subtypes, and constructed CMLS-related prognostic model features, which are important for accurate management and effective treatment of STAD.

PMID:39885324 | DOI:10.1038/s41598-025-87444-3

❌