Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
ConSensus: Multi-Agent Collaboration for Multimodal Sensing
arXiv:2601.06453v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly grounded in sensor data to perceive and reason about human physiology and the physical world. However, accurately interpreting heterogeneous multimodal sensor data remains a fundamental challenge. We show that a single monolithic LLM often fails to reason coherently across modalities, leading to incomplete interpretations and prior-knowledge bias. We introduce ConSensus, a training-free multi-agent col
-
cs.AI, q-bio.NC updates on arXiv.org
-
SafePro: Evaluating the Safety of Professional-Level AI Agents
arXiv:2601.06663v1 Announce Type: new Abstract: Large language model-based agents are rapidly evolving from simple conversational assistants into autonomous systems capable of performing complex, professional-level tasks in various domains. While these advancements promise significant productivity gains, they also introduce critical safety risks that remain under-explored. Existing safety evaluations primarily focus on simple, daily assistance tasks, failing to capture the intricate decision-ma
SafePro: Evaluating the Safety of Professional-Level AI Agents
-
cs.AI, q-bio.NC updates on arXiv.org
-
Puzzle it Out: Local-to-Global World Model for Offline Multi-Agent Reinforcement Learning
arXiv:2601.07463v1 Announce Type: new Abstract: Offline multi-agent reinforcement learning (MARL) aims to solve cooperative decision-making problems in multi-agent systems using pre-collected datasets. Existing offline MARL methods primarily constrain training within the dataset distribution, resulting in overly conservative policies that struggle to generalize beyond the support of the data. While model-based approaches offer a promising solution by expanding the original dataset with syntheti
Puzzle it Out: Local-to-Global World Model for Offline Multi-Agent Reinforcement Learning
-
cs.AI, q-bio.NC updates on arXiv.org
-
Why Slop Matters
arXiv:2601.06060v1 Announce Type: cross Abstract: AI-generated "slop" is often seen as digital pollution. We argue that this dismissal of the topic risks missing important aspects of AI Slop that deserve rigorous study. AI Slop serves a social function: it offers a supply-side solution to a variety of problems in cultural and economic demand - that, collectively, people want more content than humans can supply. We also argue that AI Slop is not mere digital detritus but has its own aesthetic va
Why Slop Matters
-
cs.AI, q-bio.NC updates on arXiv.org
-
PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis
arXiv:2601.07344v1 Announce Type: cross Abstract: Recent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of real-world clinical diagnostics, which involve heterogeneous inputs and require ongoing contextual understanding during patient-physician interactions. To bridge this gap, we introduce PulseMind, a new family of multi-modal diagnostic models that integrates a systematica
PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis
-
cs.AI, q-bio.NC updates on arXiv.org
-
FairMedQA: Benchmarking Bias in Large Language Models for Medical Question Answering
arXiv:2505.19562v2 Announce Type: replace Abstract: Large language models (LLMs) are approaching expert-level performance in medical question answering (QA), demonstrating strong potential to improve public healthcare. However, underlying biases related to sensitive attributes such as sex and race pose life-critical risks. The extent to which such sensitive attributes affect diagnosis remains an open question and requires comprehensive empirical investigation. Additionally, even the latest Coun
FairMedQA: Benchmarking Bias in Large Language Models for Medical Question Answering
-
cs.AI, q-bio.NC updates on arXiv.org
-
Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
arXiv:2512.20387v3 Announce Type: replace Abstract: We propose a Vision-Language Simulation Model (VLSM) that unifies visual and textual understanding to synthesize executable FlexScript from layout sketches and natural-language prompts, enabling cross-modal reasoning for industrial simulation systems. To support this new paradigm, the study constructs the first large-scale dataset for generative digital twins, comprising over 120,000 prompt-sketch-code triplets that enable multimodal learning
Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
-
Nature Medicine
-
Interpretable inflammation landscape of circulating immune cells
Nature Medicine, Published online: 12 January 2026; doi:10.1038/s41591-025-04126-3Including data from 1,047 patients across 19 inflammatory diseases, a new atlas presents a comprehensive model of inflammation in circulating immune cells.
Interpretable inflammation landscape of circulating immune cells
Nature Medicine, Published online: 12 January 2026; doi:10.1038/s41591-025-04126-3
Including data from 1,047 patients across 19 inflammatory diseases, a new atlas presents a comprehensive model of inflammation in circulating immune cells.-
cs.AI, q-bio.NC updates on arXiv.org
-
Safety Not Found (404): Hidden Risks of LLM-Based Robotics Decision Making
arXiv:2601.05529v1 Announce Type: new Abstract: One mistake by an AI system in a safety-critical setting can cost lives. As Large Language Models (LLMs) become integral to robotics decision-making, the physical dimension of risk grows; a single wrong instruction can directly endanger human safety. This paper addresses the urgent need to systematically evaluate LLM performance in scenarios where even minor errors are catastrophic. Through a qualitative evaluation of a fire evacuation scenario, w
Safety Not Found (404): Hidden Risks of LLM-Based Robotics Decision Making
-
Nature Medicine
-
BCMA-directed mRNA CAR-T cell therapy for myasthenia gravis: exploratory biomarker analysis of a placebo-controlled phase 2b trial
Nature Medicine, Published online: 09 January 2026; doi:10.1038/s41591-025-04170-zAnalysis of a placebo-controlled trial of a BCMA-targeting CAR-T cell therapy in patients with myasthenia gravis shows that CAR-T cell infusion selectively remodels the systemic immune environment, with elimination of BCMA-high plasma cells and activated plasmacytoid dendritic cells and changes in the autoreactive B cell repertoire.
BCMA-directed mRNA CAR-T cell therapy for myasthenia gravis: exploratory biomarker analysis of a placebo-controlled phase 2b trial
Nature Medicine, Published online: 09 January 2026; doi:10.1038/s41591-025-04170-z
Analysis of a placebo-controlled trial of a BCMA-targeting CAR-T cell therapy in patients with myasthenia gravis shows that CAR-T cell infusion selectively remodels the systemic immune environment, with elimination of BCMA-high plasma cells and activated plasmacytoid dendritic cells and changes in the autoreactive B cell repertoire.-
cs.AI, q-bio.NC updates on arXiv.org
-
Surface-based Molecular Design with Multi-modal Flow Matching
arXiv:2601.04506v1 Announce Type: cross Abstract: Therapeutic peptides show promise in targeting previously undruggable binding sites, with recent advancements in deep generative models enabling full-atom peptide co-design for specific protein receptors. However, the critical role of molecular surfaces in protein-protein interactions (PPIs) has been underexplored. To bridge this gap, we propose an omni-design peptides generation paradigm, called SurfFlow, a novel surface-based generative algori
Surface-based Molecular Design with Multi-modal Flow Matching
-
cs.AI, q-bio.NC updates on arXiv.org
-
Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework
arXiv:2601.04790v1 Announce Type: cross Abstract: Multi-agent systems utilizing large language models often assign authoritative roles to improve performance, yet the impact of authority bias on agent interactions remains underexplored. We present the first systematic analysis of role-based authority bias in free-form multi-agent evaluation using ChatEval. Applying French and Raven's power-based theory, we classify authoritative roles into legitimate, referent, and expert types and analyze thei
Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework
-
cs.AI, q-bio.NC updates on arXiv.org
-
Atlas 2 -- Foundation models for clinical deployment
arXiv:2601.05148v1 Announce Type: cross Abstract: Pathology foundation models substantially advanced the possibilities in computational pathology -- yet tradeoffs in terms of performance, robustness, and computational requirements remained, which limited their clinical deployment. In this report, we present Atlas 2, Atlas 2-B, and Atlas 2-S, three pathology vision foundation models which bridge these shortcomings by showing state-of-the-art performance in prediction performance, robustness, and
Atlas 2 -- Foundation models for clinical deployment
-
cs.AI, q-bio.NC updates on arXiv.org
-
PsychEval: A Multi-Session and Multi-Therapy Benchmark for High-Realism AI Psychological Counselor
arXiv:2601.01802v3 Announce Type: replace Abstract: To develop a reliable AI for psychological assessment, we introduce \texttt{PsychEval}, a multi-session, multi-therapy, and highly realistic benchmark designed to address three key challenges: \textbf{1) Can we train a highly realistic AI counselor?} Realistic counseling is a longitudinal task requiring sustained memory and dynamic goal tracking. We propose a multi-session benchmark (spanning 6-10 sessions across three distinct stages) that de
PsychEval: A Multi-Session and Multi-Therapy Benchmark for High-Realism AI Psychological Counselor
-
Journal of Medical Internet Research
-
A Web-Based Cancer Prevention Intervention for Rural Emerging Adults: Mixed Methods Development and Pilot-Testing Study
Background: The rapid growth of user-generated web-based health information increases the complexity of cancer information seeking. One promising strategy for promoting high-quality cancer information consumption is through targeted interventions that are intentionally designed to reach individuals in the web-based spaces they occupy. However, there is a paucity of evidence-based information on the best strategies for designing and implementing web-based health behavior change interventions to i
A Web-Based Cancer Prevention Intervention for Rural Emerging Adults: Mixed Methods Development and Pilot-Testing Study
-
npj Digital Medicine
-
An autonomous agentic workflow for clinical detection of cognitive concerns using large language models
npj Digital Medicine, Published online: 07 January 2026; doi:10.1038/s41746-025-02324-4An autonomous agentic workflow for clinical detection of cognitive concerns using large language models
An autonomous agentic workflow for clinical detection of cognitive concerns using large language models
npj Digital Medicine, Published online: 07 January 2026; doi:10.1038/s41746-025-02324-4
An autonomous agentic workflow for clinical detection of cognitive concerns using large language models-
Omics In Lung
-
Adaptive therapy for perioperative non-small cell lung cancer: strategies guided by dynamic minimal residual disease adjustment
Transl Oncol. 2026 Jan 6;64:102660. doi: 10.1016/j.tranon.2025.102660. Online ahead of print.ABSTRACTLung cancer remains the leading cause of cancer incidence and mortality worldwide, with non-small cell lung cancer (NSCLC) accounting for about 85% of cases. The low rate of early diagnosis and the high rate of occult metastases limit the survival benefits of conventional treatments. The current TNM staging system fails to fully reflect tumor heterogeneity or the dynamic molecular evolution of th
Adaptive therapy for perioperative non-small cell lung cancer: strategies guided by dynamic minimal residual disease adjustment
Transl Oncol. 2026 Jan 6;64:102660. doi: 10.1016/j.tranon.2025.102660. Online ahead of print.
ABSTRACT
Lung cancer remains the leading cause of cancer incidence and mortality worldwide, with non-small cell lung cancer (NSCLC) accounting for about 85% of cases. The low rate of early diagnosis and the high rate of occult metastases limit the survival benefits of conventional treatments. The current TNM staging system fails to fully reflect tumor heterogeneity or the dynamic molecular evolution of the disease, thus affecting the prediction of recurrence and the prognostic stratification. Some recent advances in minimal residual disease (MRD) detection, such as ultra-sensitive liquid biopsy technologies, have largely overcome the limitations of traditional imaging and offered a transformative approach for continuous, precision-based management of lung cancer. This review systematically summarized the technological evolution of MRD detection and highlighted its clinical significance in guiding adaptive therapy for NSCLC, including treatment escalation, de-escalation, and the emerging concept of precision-guided drug holidays. Moreover, the authors comprehensively discussed the "Four-Dimensional TNMB Staging System," which incorporates continuous molecular monitoring to address the static limitations of conventional staging and enhance the accuracy of prognostic stratification. Although ongoing challenges, such as the lack of standardized interpretation criteria and limited detection sensitivity, the combinations with the third-generation liquid biopsy platforms, multi-omics analyses, and multi-center prospective validation studies are expected to advance the clinical implementation of MRD-guided strategies. The paradigm change will enable the transition of NSCLC management from conventional standardized models to a precision-guided, closed-loop system of "monitoring-intervention-remonitoring," establishing a solid theoretical and practical foundation for comprehensive, molecularly driven management strategies.
PMID:41496417 | DOI:10.1016/j.tranon.2025.102660
-
Journal of Medical Internet Research
-
Establishment and Optimization of a Patient-Reported Outcome–Based Electronic-Diary for Symptoms Evaluation in Patients With Gastroesophageal Reflux Disorder: Prospective Cohort Study
Background: Gastroesophageal reflux disease (GERD) symptoms significantly affect patients’ quality of life. Patient-reported outcome (PRO) instruments for symptoms measurement in GERD patients is advocated by regulatory authority. Current tools for GERD symptoms evaluation are limited and the results can be biased by the recall bias. To better characterize the GERD symptoms, an e-diary was developed for daily GERD symptom monitoring. Objective: To build up and optimize a PRO-based e-diary, and t
Establishment and Optimization of a Patient-Reported Outcome–Based Electronic-Diary for Symptoms Evaluation in Patients With Gastroesophageal Reflux Disorder: Prospective Cohort Study
-
cs.AI, q-bio.NC updates on arXiv.org
-
OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment
arXiv:2601.01576v1 Announce Type: cross Abstract: Evaluating novelty is critical yet challenging in peer review, as reviewers must assess submissions against a vast, rapidly evolving literature. This report presents OpenNovelty, an LLM-powered agentic system for transparent, evidence-based novelty analysis. The system operates through four phases: (1) extracting the core task and contribution claims to generate retrieval queries; (2) retrieving relevant prior work based on extracted queries via
OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment
-
cs.AI, q-bio.NC updates on arXiv.org
-
JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models
arXiv:2601.01627v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed in healthcare field, it becomes essential to carefully evaluate their medical safety before clinical use. However, existing safety benchmarks remain predominantly English-centric, and test with only single-turn prompts despite multi-turn clinical consultations. To address these gaps, we introduce JMedEthicBench, the first multi-turn conversational benchmark for evaluating medical safety o