❌

Normal view

XR-DT: Extended Reality-Enhanced Digital Twin for Agentic Mobile Robots

arXiv:2512.05270v1 Announce Type: cross Abstract: As mobile robots increasingly operate alongside humans in shared workspaces, ensuring safe, efficient, and interpretable Human-Robot Interaction (HRI) has become a pressing challenge. While substantial progress has been devoted to human behavior prediction, limited attention has been paid to how humans perceive, interpret, and trust robots' inferences, impeding deployment in safety-critical and socially embedded environments. This paper presents XR-DT, an eXtended Reality-enhanced Digital Twin framework for agentic mobile robots, that bridges physical and virtual spaces to enable bi-directional understanding between humans and robots. Our hierarchical XR-DT architecture integrates virtual-, augmented-, and mixed-reality layers, fusing real-time sensor data, simulated environments in the Unity game engine, and human feedback captured through wearable AR devices. Within this framework, we design an agentic mobile robot system with a unified diffusion policy for context-aware task adaptation. We further propose a chain-of-thought prompting mechanism that allows multimodal large language models to reason over human instructions and environmental context, while leveraging an AutoGen-based multi-agent coordination layer to enhance robustness and collaboration in dynamic tasks. Initial experimental results demonstrate accurate human and robot trajectory prediction, validating the XR-DT framework's effectiveness in HRI tasks. By embedding human intention, environmental dynamics, and robot cognition into the XR-DT framework, our system enables interpretable, trustworthy, and adaptive HRI.

M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG

arXiv:2512.05959v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved strong performance in visual question answering (VQA), yet they remain constrained by static training data. Retrieval-Augmented Generation (RAG) mitigates this limitation by enabling access to up-to-date, culturally grounded, and multilingual information; however, multilingual multimodal RAG remains largely underexplored. We introduce M4-RAG, a massive-scale benchmark covering 42 languages and 56 regional dialects and registers, comprising over 80,000 culturally diverse image-question pairs for evaluating retrieval-augmented VQA across languages and modalities. To balance realism with reproducibility, we build a controlled retrieval environment containing millions of carefully curated multilingual documents relevant to the query domains, approximating real-world retrieval conditions while ensuring consistent experimentation. Our systematic evaluation reveals that although RAG consistently benefits smaller VLMs, it fails to scale to larger models and often even degrades their performance, exposing a critical mismatch between model size and current retrieval effectiveness. M4-RAG provides a foundation for advancing next-generation RAG systems capable of reasoning seamlessly across languages, modalities, and cultural contexts.

The AI Productivity Index (APEX)

arXiv:2509.25721v4 Announce Type: replace-cross Abstract: We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases per job (n = 400 total) and updates to the grading methodology. We present a new leaderboard, where GPT5 (Thinking = High) remains the top performing model with a score of 67.0%. APEX-v1-extended shows that frontier models still have substantial limitations when performing typical professional tasks. To support further research, we are open sourcing n = 25 non-benchmark example cases per role (n = 100 total) along with our evaluation harness.

Concept-Guided Backdoor Attack on Vision Language Models

arXiv:2512.00713v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have achieved impressive progress in multimodal text generation, yet their rapid adoption raises increasing concerns about security vulnerabilities. Existing backdoor attacks against VLMs primarily rely on explicit pixel-level triggers or imperceptible perturbations injected into images. While effective, these approaches reduce stealthiness and remain vulnerable to image-based defenses. We introduce concept-guided backdoor attacks, a new paradigm that operates at the semantic concept level rather than on raw pixels. We propose two different attacks. The first, Concept-Thresholding Poisoning (CTP), uses explicit concepts in natural images as triggers: only samples containing the target concept are poisoned, causing the model to behave normally in all other cases but consistently inject malicious outputs whenever the concept appears. The second, CBL-Guided Unseen Backdoor (CGUB), leverages a Concept Bottleneck Model (CBM) during training to intervene on internal concept activations, while discarding the CBM branch at inference time to keep the VLM unchanged. This design enables systematic replacement of a targeted label in generated text (for example, replacing "cat" with "dog"), even when the replacement behavior never appears in the training data. Experiments across multiple VLM architectures and datasets show that both CTP and CGUB achieve high attack success rates while maintaining moderate impact on clean-task performance. These findings highlight concept-level vulnerabilities as a critical new attack surface for VLMs.

Critical Appraisal Tools for Evaluating Artificial Intelligence in Clinical Studies: Scoping Review

Background: Health research that uses predictive and/or generative AI is rapidly growing. Just as in traditional clinical studies, the way in which AI studies are conducted can introduce systematic errors. Transmission of this AI evidence into clinical practice and research needs critical appraisal tools for clinical decision makers and researchers. Objective: To identify existing tools for critical appraisal of clinical studies that use artificial intelligence (AI) and examine the concepts and domains these tools explore. Methods: Inclusion criteria in PCC framework P: (population) Artificial intelligence clinical studies. C (Concept): tools for critical appraisal and associated constructs such as: quality, reporting, validity, risk of bias, and applicability. C (context): in clinical practice context. In addition, bias classification and Chatbot assessment studies were included. We searched in medical and engineering databases (MEDLINE, EMBASE, CINAHL, PsycINFO and IEEE). We included clinical primary research with tools for critical appraisal. Classic reviews and systematic reviews were included in first phase of screening. They were excluded in the secondary phase, after identifying new tools by forward snowballing. We excluded non-human, computer and mathematical research, and letters, opinion papers and editorials. We used Rayyan for screening. Data extraction was done by two observers and discrepancies were solved by discussion. The protocol was previously registered in OSF (https://doi.org/10.17605/OSF.IO/ETYDS). We adhered to the PRISMA extension for Scoping reviews and to the PRISMA-Search extension for Reporting Literature in Systematic Reviews. Results: We retrieved 4393 unique records for screening. After excluding 3803 records, 119 were selected for full-text screening. From these, 59 were excluded. After inclusion of 10 studies via other methods, a total of 70 records were finally included. 46 of them were reporting guidelines (15 tools for critical appraisal, 2 for quality of study and 2 for risk of bias). Nine papers ware focused on bias classification or mitigation. We found 15 Chatbots assessment studies or systematic reviews of Chatbots studies (6 and 9 respectively) which are a very heterogeneous group. Conclusions: The results picture a landscape of the evidence tools where reporting tools predominate, followed by critical appraisal and risk of bias tools, and few tools for risk of bias. The mismatch of bias in AI and epidemiology should be considered for critical appraisal, especially regarding fairness and the mitigation bias in the AI. Finally, Chatbot assessment studies is a vast and evolving field in which progress in design, reporting and critical appraisal is necessary and urgent. Clinical Trial: https://doi.org/10.17605/OSF.IO/ETYDS

Exploring a Digital Health Solution to Collect and Manage Health-Related Needs for Patients Who Undergo Complex Surgery: Mixed Methods Study

Background: Patients who undergo complex surgery (e.g., esophagectomy, liver resection) often experience substantial burden of health-related needs (medical, social, and behavioral health). A closed loop digital solution could facilitate the collection and resolution of health-related needs by care team members for patients who undergo complex surgery. A digital solution may facilitate adherence to a clear treatment plan and concomitantly reduce surgical complications and readmissions associated with unmet health-related needs, which remain persistent challenges across health care settings. Objective: To establish problems and gaps in the collection, integration, and management of health-related needs and identify a set of user specifications for a digital solution to collect and manage health-related needs, specifically medical, social, and behavioral needs for patients who undergo complex surgery. Methods: We applied the Double Diamond Framework and organized the study into two sequential phases: (1) qualitative methods to discover patients’ and care team members’ perspectives on health-related needs; (2) participatory design sessions to gain feedback and sentiment about ideal features of a digital solution. Both phases were conducted between December 2023 and March 2025. We supplemented both phases with analysis of electronic health record (EHR) data for patients who underwent complex surgery at our academic medical center (AMC). Results: Extensive themes emerged from interviews with patients (n=20) and care team members (n=24), capturing their health-related and surgical experiences as well as desired features for a proposed digital solution. Our swim lane diagram demonstrated four critical gaps in workflow: (1) heterogeneity in the approach to screening, monitoring, and managing health-related needs; (2) patients felt uncomfortable reporting health-related needs, particularly behavioral and social needs, to their care team; (3) lack of access to referral resources to resolve needs; and (4) the need for a closed loop intervention for patients and care team members. A subset of participants from Phase 1 (n=5 patients and n=9 care team members) provided feedback on preferred features, drawing from digital tools currently available in the EHR at our AMC. Among four existing EHR tools tested, there were notable variations in how patients and care team members felt about their potential use. Participants also provided extensive feedback for preferred components (e.g., goals and active plans) that should be available in an existing or custom digital solution to manage health-related needs. Findings from the qualitative interviews and design sessions were corroborated with EHR documentation. Conclusions: Digital solutions could provide a streamlined approach for collection and management of health-related needs in surgery, with the goal of addressing unmet needs and improving patient activation. This approach is critical to ensure patients, especially patients who undergo complex surgery, have positive health outcomes. We identified preferences for specific features in a proposed digital solution based on our systematic assessment that will inform future work.

Molecular stratification of esophageal adenocarcinoma: implications for prognosis and treatment strategy

Oncogene, Published online: 08 December 2025; doi:10.1038/s41388-025-03650-3

Molecular stratification of esophageal adenocarcinoma: implications for prognosis and treatment strategy

AI-driven transfer learning and classical molecular dynamics for strategic therapeutic repurposing and rational design of antiviral peptides targeting monkeypox virus DNA polymerase

Comput Biol Med. 2025 Dec 7;200:111372. doi: 10.1016/j.compbiomed.2025.111372. Online ahead of print.

ABSTRACT

The emergence of monkeypox virus (MPXV) as a global health threat has necessitated the rapid identification of novel antiviral therapeutics. Currently, no FDA-approved drugs are specifically designed against the disease. We used an in-house deep learning pharmacophore model for screening a library of 1974 FDA-approved drugs targeting the active site of MPXV DNA polymerase. Three drugs exhibited the strongest binding affinities, outperforming the control drug, Cidofovir diphosphate, and forming stable interactions with key active site residues. Among them, Paromomycin emerged as the most favourable drug, demonstrating stable, persistent, and adaptable interactions in molecular dynamics simulation. In parallel, we developed a novel automated peptide-generating AI pipeline that integrates active-site residues with knowledge-guided amino acid selection to generate and evaluate synthetic peptides. Cysteine-Phenylalanine-Cysteine (CFC), together with a panel of candidates, emerged through rational balancing of physicochemical properties and drug-likeness for accelerated therapeutic discovery. Synthetic peptides were evaluated to further understand the binding efficacies with DNA polymerase. CFC peptide demonstrated strong binding affinity (-8.08 kcal/mol) through stable interactions with key catalytic residues ASP549, ARG634 and LYS661, while MMGBSA analysis confirmed favourable binding energy (-33.02 kcal/mol). Consistent results in MD simulations indicate functional binding without destabilisation. Although ADMET predictions for CFC revealed limitations in permeability and oral bioavailability, its favourable binding profile and reduced predicted toxicity support its potential as a novel antiviral lead.

PMID:41360016 | DOI:10.1016/j.compbiomed.2025.111372

❌