❌

Reading view

Intervention in Health Misinformation Using Large Language Models for Automated Detection, Thematic Analysis, and Inoculation: Case Study on COVID-19

Background: The rapid growth of social media as an information channel has enabled the swift spread of inaccurate or false health information, significantly impacting public health. This widespread dissemination of misinformation has caused confusion, eroded trust in health authorities, led to noncompliance with health guidelines, and encouraged risky health behaviors. Understanding the dynamics of misinformation on social media is essential for devising effective public health communication strategies. Objective: This study aims to present a comprehensive and automated approach that leverages Large Language Models (LLMs) and Machine Learning (ML) techniques to detect misinformation on social media, uncover the underlying causes and themes, and generate refutation arguments, facilitating control of its spread and promoting public health outcomes by inoculating people against health misinformation. Methods: We use two datasets to train three LLMs, namely BERT, T5, and GPT-2, to classify documents into two categories: misinformation and non-misinformation. Additionally, we employ a separate dataset to identify misinformation topics. To analyze these topics, we applied three topic modeling algorithms—Latent Dirichlet Allocation (LDA), Top2Vec, and BERTopic—and selected the optimal model based on performance evaluated across three metrics. Using a prompting approach, we extract sentence-level representations for the topics to uncover their underlying themes. Finally, we design a prompt text capable of identifying misinformation themes effectively. Codes are available at https://github.com/SamiraMalek/MDIP-MDIS. Results: The trained BERT model demonstrated exceptional performance, achieving 98% accuracy in classifying misinformation and non-misinformation, with a 44% reduction in false positive rates for AI-generated misinformation. Among the three topic modeling approaches employed, BERTopic outperformed the others, achieving the highest metrics with a Coherence Value (CV) of 0.41, Normalized Pointwise Mutual Information (NPMI) of -0.086, and Inverted RBO (IRBO) of 0.99. To address the issue of unclassified documents, we developed an algorithm to assign each document to its closest topic. Additionally, we proposed a novel method using prompt engineering to generate sentence-level representations for each topic, achieving a 99.6% approval rate as "appropriate" or "somewhat appropriate" by three independent raters. We further designed a prompt text to identify themes of misinformation topics and developed another prompt capable of detecting misinformation themes with 82% accuracy. Conclusions: This study presents a comprehensive and automated approach to addressing health misinformation on social media using advanced machine learning and natural language processing techniques. By leveraging large language models (LLMs) and prompt engineering, the system effectively detects misinformation, identifies underlying themes, and provides explanatory responses to combat its spread. The proposed method was tested on an English-language COVID-19–related dataset and has not been evaluated on real-world online social media data; the experiments were conducted offline.
  •  

Digital Twin based Automatic Reconfiguration of Robotic Systems in Smart Environments

arXiv:2511.00094v2 Announce Type: replace-cross Abstract: Robotic systems have become integral to smart environments, enabling applications ranging from urban surveillance and automated agriculture to industrial automation. However, their effective operation in dynamic settings - such as smart cities and precision farming - is challenged by continuously evolving topographies and environmental conditions. Traditional control systems often struggle to adapt quickly, leading to inefficiencies or operational failures. To address this limitation, we propose a novel framework for autonomous and dynamic reconfiguration of robotic controllers using Digital Twin technology. Our approach leverages a virtual replica of the robot's operational environment to simulate and optimize movement trajectories in response to real-world changes. By recalculating paths and control parameters in the Digital Twin and deploying the updated code to the physical robot, our method ensures rapid and reliable adaptation without manual intervention. This work advances the integration of Digital Twins in robotics, offering a scalable solution for enhancing autonomy in smart, dynamic environments.
  •  

Computational network models for forecasting and control of mental health trajectories in digital applications

npj Digital Medicine, Published online: 30 December 2025; doi:10.1038/s41746-025-02252-3

Computational network models for forecasting and control of mental health trajectories in digital applications
  •  

Metabolic signatures in gastroenteropancreatic neuroendocrine neoplasms: unraveling diagnostic and prognostic insights

Front Endocrinol (Lausanne). 2025 Dec 11;16:1676021. doi: 10.3389/fendo.2025.1676021. eCollection 2025.

ABSTRACT

Gastroenteropancreatic neuroendocrine neoplasms (GEP-NENs) are a heterogeneous group of tumors characterized by diverse biological behaviors and variable clinical outcomes. Recent advances have highlighted the important role of metabolic reprogramming in tumorigenesis, progression, and therapeutic resistance in GEP-NENs. In this review, we synthesize the current evidence on metabolic biomarkers and altered metabolic pathways-particularly those involving glucose, lipid, and amino acid metabolism. Key biomarkers such as GLUT-1, FASN, and enzymes involved in ferroptosis, cholesterol biosynthesis, and amino acid catabolism demonstrate strong associations with tumor aggressiveness, hypoxia, and mTOR signaling. Moreover, metabolomic profiling and functional studies suggest that metabolic markers may inform prognosis and predict response to targeted therapies such as Everolimus. Although promising, the clinical translation of these markers is still limited and requires further validation in large, subtype-specific cohorts. Our findings highlight the importance of integrating metabolic profiling into the diagnostic and therapeutic landscape of GEP-NENs. Future research should prioritize biomarker standardization, multi-omics integration, and the development of metabolism-based therapeutic strategies tailored to tumor subtype and differentiation grade.

PMID:41458541 | PMC:PMC12738315 | DOI:10.3389/fendo.2025.1676021

  •  

Evaluating Peer Online Forums to Support Health: Ethical and Practical Challenges

Many people use peer online forums to seek support for health-related problems. More research is needed to understand the impacts of forum use, and how these are generated. However, there are significant ethical and practical challenges with the methods available to do the required research. We examine the key challenges associated with conducting each of the most commonly used online data collection methods: surveys, interviews, forum post analysis; and triangulation of these methods. Based on our learning from the Improving Peer Online Forums (iPOF) study, an inter-disciplinary realist informed mixed methods evaluation of peer online forums, we outline strategies that can be used to address key issues pertaining to assessing important outcomes, facilitating participation, validating participants (users who consent to take part in one or more parts of the study), protecting anonymity, gaining consent, managing risk, multi-stakeholder engagement, and triangulation. We share this learning to support researchers, reviewers, and ethics committees faced with deciding how best to address these challenges. We highlight the need for ongoing open, transparent discussion to ensure the research field keeps pace with evolving technology design and societal attitudes to online data use.
  •  

XTC, A Research Platform for Optimizing AI Workload Operators

arXiv:2512.16512v1 Announce Type: cross Abstract: Achieving high efficiency on AI operators demands precise control over computation and data movement. However, existing scheduling languages are locked into specific compiler ecosystems, preventing fair comparison, reuse, and evaluation across frameworks. No unified interface currently decouples scheduling specification from code generation and measurement. We introduce XTC, a platform that unifies scheduling and performance evaluation across compilers. With its common API and reproducible measurement framework, XTC enables portable experimentation and accelerates research on optimization strategies.
  •  

First, do NOHARM: towards clinically safe large language models

arXiv:2512.01241v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized. We present NOHARM (Numerous Options Harm Assessment for Risk in Medicine), a benchmark using 100 real primary care-to-specialist consultation cases to measure frequency and severity of harm from LLM-generated medical recommendations. NOHARM covers 10 specialties, with 12,747 expert annotations for 4,249 clinical management options. Across 31 LLMs, potential for severe harm from LLM recommendations occurs in up to 22.2% (95% CI 21.6-22.8%) of cases, with harm of omission accounting for 76.6% (95% CI 76.4-76.8%) of errors. Safety performance is only moderately correlated (r = 0.61-0.64) with existing AI and medical knowledge benchmarks. The best models outperform generalist physicians on safety (mean difference 9.7%, 95% CI 7.0-12.5%), and a diverse multi-agent approach improves safety compared to solo models (mean difference 8.0%, 95% CI 4.0-12.1%). Therefore, despite strong performance on existing evaluations, widely used AI models can produce severely harmful medical advice at nontrivial rates, underscoring clinical safety as a distinct performance dimension necessitating explicit measurement.
  •  

A Multicenter Benchmark of Multiple Instance Learning Models for Lymphoma Subtyping from HE-stained Whole Slide Images

arXiv:2512.14640v1 Announce Type: cross Abstract: Timely and accurate lymphoma diagnosis is essential for guiding cancer treatment. Standard diagnostic practice combines hematoxylin and eosin (HE)-stained whole slide images with immunohistochemistry, flow cytometry, and molecular genetic tests to determine lymphoma subtypes, a process requiring costly equipment, skilled personnel, and causing treatment delays. Deep learning methods could assist pathologists by extracting diagnostic information from routinely available HE-stained slides, yet comprehensive benchmarks for lymphoma subtyping on multicenter data are lacking. In this work, we present the first multicenter lymphoma benchmarking dataset covering four common lymphoma subtypes and healthy control tissue. We systematically evaluate five publicly available pathology foundation models (H-optimus-1, H0-mini, Virchow2, UNI2, Titan) combined with attention-based (AB-MIL) and transformer-based (TransMIL) multiple instance learning aggregators across three magnifications (10x, 20x, 40x). On in-distribution test sets, models achieve multiclass balanced accuracies exceeding 80% across all magnifications, with all foundation models performing similarly and both aggregation methods showing comparable results. The magnification study reveals that 40x resolution is sufficient, with no performance gains from higher resolutions or cross-magnification aggregation. However, on out-of-distribution test sets, performance drops substantially to around 60%, highlighting significant generalization challenges. To advance the field, larger multicenter studies covering additional rare lymphoma subtypes are needed. We provide an automated benchmarking pipeline to facilitate such future research.
  •  

Generating Reliable Synthetic Clinical Trial Data: The Role of Hyperparameter Optimization and Domain Constraints

arXiv:2505.05019v2 Announce Type: replace-cross Abstract: The generation of synthetic clinical trial data offers a promising approach to mitigating privacy concerns and data accessibility limitations in medical research. However, ensuring that synthetic datasets maintain high fidelity, utility, and adherence to domain-specific constraints remains a key challenge. While hyperparameter optimization (HPO) improves generative model performance, the effectiveness of different optimization strategies for synthetic clinical data remains unclear. This study systematically evaluates four HPO objectives across nine generative models, comparing single-metric to compound metric optimization. Our results demonstrate that HPO consistently improves synthetic data quality, with Tab DDPM achieving the largest relative gains, followed by TVAE (60%), CTGAN (39%), and CTAB-GAN+ (38%). Compound metric optimization outperformed single-metric objectives, producing more generalizable synthetic datasets. Despite improving overall quality, HPO alone fails to prevent violations of essential clinical survival constraints. Preprocessing and postprocessing played a crucial role in reducing these violations, as models lacking robust processing steps produced invalid data in up to 61% of cases. These findings underscore the necessity of integrating explicit domain knowledge alongside HPO to generate high-quality synthetic datasets. Our study provides actionable recommendations for improving synthetic data generation, with future work needed to refine metric selection and validate findings on larger datasets.
  •  

Data Visualization Support for Interdisciplinary Team Treatment Planning in Clinical Oncology: Scoping Review

Background: Complex and expanding datasets in clinical oncology applications require flexible and interactive visualization of patient data to provide physicians and other medical professionals with maximum amount of information. In particular, interdisciplinary tumor conferences profit from customized tools to integrate, link, and visualize relevant data from all professions involved. Objective: Our objective was to identify and present currently available data visualization tools for tumor boards and related areas. We wanted to provide an overview of not only the digital tools currently used in tumor board settings but also of the data they include, their respective visualization solutions, and their integration into hospital processes. Methods: This scoping review was based on the scoping study framework by Arksey and O’Malley and attempted to answer the following research question: “What are the key features of data visualization solutions used in molecular and organ tumor boards, and how are these elements integrated and used within the clinical setting?” The following electronic databases were searched for articles: PubMed, Web of Science, and Scopus. Articles were deemed eligible if published in English in the last 10 years. Eligible articles were first deduplicated, followed by screening of titles and abstracts. Full-text screening was then conducted to decide on article selection. All included articles were analyzed using a data extraction template. The template included a variety of meta-information, as well as specific fields aiming to answer the research question. Results: The review process started with 2049 articles, of which 1014 (49.49%) were included in the title and abstract screening. A total of 5.47% (112/2049) of the publications were eligible for full-text screening, leading to 2.93% (60/2049) of the publications being eligible for final inclusion. They covered 49 distinct visualization tools and applications. We discovered a variety of innovative visualization solutions, most often driven by the complexity of omics data, represented in 96% (47/49) of the tools. Tables remained the most used tool for the visualization of data types described in the articles. Approximately one-third of the identified tools (16/49, 33%) were systematically evaluated in some form. For most discovered tools (37/49, 76%), there was no documentation of implementation into the clinical routine. A significant number of applications (21/49, 43%) were available through open-source access. Conclusions: There is a wide range of projects providing visualization solutions for tumor boards and clinical oncology applications. Among the few tools that have made their way into clinical routine settings, there are both commercial and academic solutions. While tables for a variety of data types remain the dominant visualization strategy, the complexity of omics data appears to be the driving force behind many visualization innovations in the domain of tumor boards. Trial Registration:
  •  

XR-DT: Extended Reality-Enhanced Digital Twin for Agentic Mobile Robots

arXiv:2512.05270v1 Announce Type: cross Abstract: As mobile robots increasingly operate alongside humans in shared workspaces, ensuring safe, efficient, and interpretable Human-Robot Interaction (HRI) has become a pressing challenge. While substantial progress has been devoted to human behavior prediction, limited attention has been paid to how humans perceive, interpret, and trust robots' inferences, impeding deployment in safety-critical and socially embedded environments. This paper presents XR-DT, an eXtended Reality-enhanced Digital Twin framework for agentic mobile robots, that bridges physical and virtual spaces to enable bi-directional understanding between humans and robots. Our hierarchical XR-DT architecture integrates virtual-, augmented-, and mixed-reality layers, fusing real-time sensor data, simulated environments in the Unity game engine, and human feedback captured through wearable AR devices. Within this framework, we design an agentic mobile robot system with a unified diffusion policy for context-aware task adaptation. We further propose a chain-of-thought prompting mechanism that allows multimodal large language models to reason over human instructions and environmental context, while leveraging an AutoGen-based multi-agent coordination layer to enhance robustness and collaboration in dynamic tasks. Initial experimental results demonstrate accurate human and robot trajectory prediction, validating the XR-DT framework's effectiveness in HRI tasks. By embedding human intention, environmental dynamics, and robot cognition into the XR-DT framework, our system enables interpretable, trustworthy, and adaptive HRI.
  •  

A Definition of AGI

arXiv:2510.18212v3 Announce Type: replace Abstract: The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quantifiable framework to address this, defining AGI as matching the cognitive versatility and proficiency of a well-educated adult. To operationalize this, we ground our methodology in Cattell-Horn-Carroll theory, the most empirically validated model of human cognition. The framework dissects general intelligence into ten core cognitive domains-including reasoning, memory, and perception-and adapts established human psychometric batteries to evaluate AI systems. Application of this framework reveals a highly "jagged" cognitive profile in contemporary models. While proficient in knowledge-intensive domains, current AI systems have critical deficits in foundational cognitive machinery, particularly long-term memory storage. The resulting AGI scores (e.g., GPT-4 at 27%, GPT-5 at 57%) concretely quantify both rapid progress and the substantial gap remaining before AGI.
  •  

Rethinking Lung Cancer Screening: AI Nodule Detection and Diagnosis Outperforms Radiologists, Leading Models, and Standards Beyond Size and Growth

arXiv:2512.00281v1 Announce Type: cross Abstract: Early detection of malignant lung nodules is critical, but its dependence on size and growth in screening inherently delays diagnosis. We present an AI system that redefines lung cancer screening by performing both detection and malignancy diagnosis directly at the nodule level on low-dose CT scans. To address limitations in dataset scale and explainability, we designed an ensemble of shallow deep learning and feature-based specialized models. Trained and evaluated on 25,709 scans with 69,449 annotated nodules, the system outperforms radiologists, Lung-RADS, and leading AI models (Sybil, Brock, Google, Kaggle). It achieves an area under the receiver operating characteristic curve (AUC) of 0.98 internally and 0.945 on an independent cohort. With 0.5 false positives per scan at 99.3\% sensitivity, it addresses key barriers to AI adoption. Critically, it outperforms radiologists across all nodule sizes and stages, excelling in stage 1 cancers, and all growth-based metrics, including the least accurate: Volume-Doubling Time. It also surpasses radiologists by up to one year in diagnosing indeterminate and slow-growing nodules.
  •  

Human Decision-making is Susceptible to AI-driven Manipulation

arXiv:2502.07663v3 Announce Type: replace Abstract: AI systems are increasingly intertwined with daily life, assisting users with various tasks and guiding decision-making. This integration introduces risks of AI-driven manipulation, where such systems may exploit users' cognitive biases and emotional vulnerabilities to steer them toward harmful outcomes. Through a randomized between-subjects experiment with 233 participants, we examined human susceptibility to such manipulation in financial (e.g., purchases) and emotional (e.g., conflict resolution) decision-making contexts. Participants interacted with one of three AI agents: a neutral agent (NA) optimizing for user benefit without explicit influence, a manipulative agent (MA) designed to covertly influence beliefs and behaviors, or a strategy-enhanced manipulative agent (SEMA) equipped with established psychological tactics, allowing it to select and apply them adaptively during interactions to reach its hidden objectives. By analyzing participants' preference ratings, we found significant susceptibility to AI-driven manipulation. Particularly across both decision-making domains, interacting with the manipulative agents significantly increased the odds of rating hidden incentives higher than optimal options (Financial, MA: OR=5.24, SEMA: OR=7.96; Emotional, MA: OR=5.52, SEMA: OR=5.71) compared to the NA group. Notably, we found no clear evidence that employing psychological strategies (SEMA) was overall more effective than simple manipulative objectives (MA) on our primary outcomes. Hence, AI-driven manipulation could become widespread even without requiring sophisticated tactics and expertise. While our findings are preliminary and derived from hypothetical, low-stakes scenarios, we highlight a critical vulnerability in human-AI interactions, emphasizing the need for ethical safeguards and regulatory frameworks to protect human autonomy.
  •  

Latent plasticity of the human pancreas across development, health, and disease

bioRxiv [Preprint]. 2025 Oct 3:2025.10.01.679230. doi: 10.1101/2025.10.01.679230.

ABSTRACT

The pancreas plays a central role in major human diseases, yet our understanding of its cellular diversity and plasticity remains incomplete. Here, we present a single-cell multiomics atlas of the human pancreas, profiling over four million cells and nuclei from 57 donors across fetal development, adult homeostasis, and type 2 diabetes (T2D). Integrating sc/snRNA-seq, snATAC-seq, VASA-seq, spatial transcriptomics (Xenium), and multiplexed proteomics (CODEX), we resolve gene expression, chromatin accessibility, and spatial organization at high resolution. We identify transcriptionally plastic centroacinar-like cells (pCACs) in adults with fetal-like features, delineate endocrine and exocrine lineage trajectories during development, and uncover HNF1A-defined beta cell epigenetic states. In T2D, we observe shifts in beta cell subtypes and altered regulatory programs. Glucose perturbation of healthy islets reveals cell-type-specific adaptation and stress responses. This atlas provides a foundational framework to understand pancreas biology and the role of cellular plasticity in regeneration and disease.

PMID:41256699 | PMC:PMC12622017 | DOI:10.1101/2025.10.01.679230

  •  

CLINB: A Climate Intelligence Benchmark for Foundational Models

arXiv:2511.11597v1 Announce Type: new Abstract: Evaluating how Large Language Models (LLMs) handle complex, specialized knowledge remains a critical challenge. We address this through the lens of climate change by introducing CLINB, a benchmark that assesses models on open-ended, grounded, multimodal question answering tasks with clear requirements for knowledge quality and evidential support. CLINB relies on a dataset of real users' questions and evaluation rubrics curated by leading climate scientists. We implement and validate a model-based evaluation process and evaluate several frontier models. Our findings reveal a critical dichotomy. Frontier models demonstrate remarkable knowledge synthesis capabilities, often exhibiting PhD-level understanding and presentation quality. They outperform "hybrid" answers curated by domain experts assisted by weaker models. However, this performance is countered by failures in grounding. The quality of evidence varies, with substantial hallucination rates for references and images. We argue that bridging this gap between knowledge synthesis and verifiable attribution is essential for the deployment of AI in scientific workflows and that reliable, interpretable benchmarks like CLINB are needed to progress towards building trustworthy AI systems.
  •  

End to End AI System for Surgical Gesture Sequence Recognition and Clinical Outcome Prediction

arXiv:2511.11899v1 Announce Type: new Abstract: Fine-grained analysis of intraoperative behavior and its impact on patient outcomes remain a longstanding challenge. We present Frame-to-Outcome (F2O), an end-to-end system that translates tissue dissection videos into gesture sequences and uncovers patterns associated with postoperative outcomes. Leveraging transformer-based spatial and temporal modeling and frame-wise classification, F2O robustly detects consecutive short (~2 seconds) gestures in the nerve-sparing step of robot-assisted radical prostatectomy (AUC: 0.80 frame-level; 0.81 video-level). F2O-derived features (gesture frequency, duration, and transitions) predicted postoperative outcomes with accuracy comparable to human annotations (0.79 vs. 0.75; overlapping 95% CI). Across 25 shared features, effect size directions were concordant with small differences (~ 0.07), and strong correlation (r = 0.96, p
  •  
❌