Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
CLINB: A Climate Intelligence Benchmark for Foundational Models
arXiv:2511.11597v1 Announce Type: new Abstract: Evaluating how Large Language Models (LLMs) handle complex, specialized knowledge remains a critical challenge. We address this through the lens of climate change by introducing CLINB, a benchmark that assesses models on open-ended, grounded, multimodal question answering tasks with clear requirements for knowledge quality and evidential support. CLINB relies on a dataset of real users' questions and evaluation rubrics curated by leading climate s
-
cs.AI, q-bio.NC updates on arXiv.org
-
End to End AI System for Surgical Gesture Sequence Recognition and Clinical Outcome Prediction
arXiv:2511.11899v1 Announce Type: new Abstract: Fine-grained analysis of intraoperative behavior and its impact on patient outcomes remain a longstanding challenge. We present Frame-to-Outcome (F2O), an end-to-end system that translates tissue dissection videos into gesture sequences and uncovers patterns associated with postoperative outcomes. Leveraging transformer-based spatial and temporal modeling and frame-wise classification, F2O robustly detects consecutive short (~2 seconds) gestures i
End to End AI System for Surgical Gesture Sequence Recognition and Clinical Outcome Prediction
-
cs.AI, q-bio.NC updates on arXiv.org
-
UpBench: A Dynamically Evolving Real-World Labor-Market Agentic Benchmark Framework Built for Human-Centric AI
arXiv:2511.12306v1 Announce Type: new Abstract: As large language model (LLM) agents increasingly undertake digital work, reliable frameworks are needed to evaluate their real-world competence, adaptability, and capacity for human collaboration. Existing benchmarks remain largely static, synthetic, or domain-limited, providing limited insight into how agents perform in dynamic, economically meaningful environments. We introduce UpBench, a dynamically evolving benchmark grounded in real jobs dra
UpBench: A Dynamically Evolving Real-World Labor-Market Agentic Benchmark Framework Built for Human-Centric AI
-
cs.AI, q-bio.NC updates on arXiv.org
-
Learning to Trust: Bayesian Adaptation to Varying Suggester Reliability in Sequential Decision Making
arXiv:2511.12378v1 Announce Type: new Abstract: Autonomous agents operating in sequential decision-making tasks under uncertainty can benefit from external action suggestions, which provide valuable guidance but inherently vary in reliability. Existing methods for incorporating such advice typically assume static and known suggester quality parameters, limiting practical deployment. We introduce a framework that dynamically learns and adapts to varying suggester reliability in partially observa
Learning to Trust: Bayesian Adaptation to Varying Suggester Reliability in Sequential Decision Making
-
cs.AI, q-bio.NC updates on arXiv.org
-
Multi-agent Self-triage System with Medical Flowcharts
arXiv:2511.12439v1 Announce Type: new Abstract: Online health resources and large language models (LLMs) are increasingly used as a first point of contact for medical decision-making, yet their reliability in healthcare remains limited by low accuracy, lack of transparency, and susceptibility to unverified information. We introduce a proof-of-concept conversational self-triage system that guides LLMs with 100 clinically validated flowcharts from the American Medical Association, providing a str
Multi-agent Self-triage System with Medical Flowcharts
-
cs.AI, q-bio.NC updates on arXiv.org
-
Conditional Diffusion Model for Multi-Agent Dynamic Task Decomposition
arXiv:2511.13137v1 Announce Type: new Abstract: Task decomposition has shown promise in complex cooperative multi-agent reinforcement learning (MARL) tasks, which enables efficient hierarchical learning for long-horizon tasks in dynamic and uncertain environments. However, learning dynamic task decomposition from scratch generally requires a large number of training samples, especially exploring the large joint action space under partial observability. In this paper, we present the Conditional
Conditional Diffusion Model for Multi-Agent Dynamic Task Decomposition
-
cs.AI, q-bio.NC updates on arXiv.org
-
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
arXiv:2511.13290v1 Announce Type: new Abstract: Humans display significant uncertainty when confronted with moral dilemmas, yet the extent of such uncertainty in machines and AI agents remains underexplored. Recent studies have confirmed the overly confident tendencies of machine-generated responses, particularly in large language models (LLMs). As these systems are increasingly embedded in ethical decision-making scenarios, it is important to understand their moral reasoning and the inherent u
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
-
cs.AI, q-bio.NC updates on arXiv.org
-
Grounded by Experience: Generative Healthcare Prediction Augmented with Hierarchical Agentic Retrieval
arXiv:2511.13293v1 Announce Type: new Abstract: Accurate healthcare prediction is critical for improving patient outcomes and reducing operational costs. Bolstered by growing reasoning capabilities, large language models (LLMs) offer a promising path to enhance healthcare predictions by drawing on their rich parametric knowledge. However, LLMs are prone to factual inaccuracies due to limitations in the reliability and coverage of their embedded knowledge. While retrieval-augmented generation (R
Grounded by Experience: Generative Healthcare Prediction Augmented with Hierarchical Agentic Retrieval
-
cs.AI, q-bio.NC updates on arXiv.org
-
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
arXiv:2511.11733v1 Announce Type: cross Abstract: Speculative decoding accelerates large language model (LLM) inference by using a lightweight draft model to propose tokens that are later verified by a stronger target model. While effective in centralized systems, its behavior in decentralized settings, where network latency often dominates compute, remains under-characterized. We present Decentralized Speculative Decoding (DSD), a plug-and-play framework for decentralized inference that turns
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
-
cs.AI, q-bio.NC updates on arXiv.org
-
Uncertainty Makes It Stable: Curiosity-Driven Quantized Mixture-of-Experts
arXiv:2511.11743v1 Announce Type: cross Abstract: Deploying deep neural networks on resource-constrained devices faces two critical challenges: maintaining accuracy under aggressive quantization while ensuring predictable inference latency. We present a curiosity-driven quantized Mixture-of-Experts framework that addresses both through Bayesian epistemic uncertainty-based routing across heterogeneous experts (BitNet ternary, 1-16 bit BitLinear, post-training quantization). Evaluated on audio cl
Uncertainty Makes It Stable: Curiosity-Driven Quantized Mixture-of-Experts
-
cs.AI, q-bio.NC updates on arXiv.org
-
Rethinking Bias in Generative Data Augmentation for Medical AI: a Frequency Recalibration Method
arXiv:2511.12301v1 Announce Type: cross Abstract: Developing Medical AI relies on large datasets and easily suffers from data scarcity. Generative data augmentation (GDA) using AI generative models offers a solution to synthesize realistic medical images. However, the bias in GDA is often underestimated in medical domains, with concerns about the risk of introducing detrimental features generated by AI and harming downstream tasks. This paper identifies the frequency misalignment between real a
Rethinking Bias in Generative Data Augmentation for Medical AI: a Frequency Recalibration Method
-
cs.AI, q-bio.NC updates on arXiv.org
-
AI Fairness Beyond Complete Demographics: Current Achievements and Future Directions
arXiv:2511.13525v1 Announce Type: cross Abstract: Fairness in artificial intelligence (AI) has become a growing concern due to discriminatory outcomes in AI-based decision-making systems. While various methods have been proposed to mitigate bias, most rely on complete demographic information, an assumption often impractical due to legal constraints and the risk of reinforcing discrimination. This survey examines fairness in AI when demographics are incomplete, addressing the gap between traditi
AI Fairness Beyond Complete Demographics: Current Achievements and Future Directions
-
cs.AI, q-bio.NC updates on arXiv.org
-
SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning
arXiv:2511.08151v2 Announce Type: replace Abstract: Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcrafted. We introduce SciAgent, a unified multi-agent system designed for generalistic scientific reasoning-the ability to adapt reasoning strategies across disciplines and difficulty levels. SciAgent organizes problem solving as a hierarchical process: a Coordinator A
SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Workflow for Full Traceability of AI Decisions
arXiv:2511.11275v2 Announce Type: replace Abstract: An ever increasing number of high-stake decisions are made or assisted by automated systems employing brittle artificial intelligence technology. There is a substantial risk that some of these decision induce harm to people, by infringing their well-being or their fundamental human rights. The state-of-the-art in AI systems makes little effort with respect to appropriate documentation of the decision process. This obstructs the ability to trac
A Workflow for Full Traceability of AI Decisions
-
cs.AI, q-bio.NC updates on arXiv.org
-
Privacy Challenges and Solutions in Retrieval-Augmented Generation-Enhanced LLMs for Healthcare Chatbots: A Review of Applications, Risks, and Future Directions
arXiv:2511.11347v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) has rapidly emerged as a transformative approach for integrating large language models into clinical and biomedical workflows. However, privacy risks, such as protected health information (PHI) exposure, remain inconsistently mitigated. This review provides a thorough analysis of the current landscape of RAG applications in healthcare, including (i) sensitive data type across clinical scenarios, (ii)
Privacy Challenges and Solutions in Retrieval-Augmented Generation-Enhanced LLMs for Healthcare Chatbots: A Review of Applications, Risks, and Future Directions
-
npj Digital Medicine
-
A large language model-based approach to quantifying the effects of social determinants in liver transplant decisions
npj Digital Medicine, Published online: 17 November 2025; doi:10.1038/s41746-025-02025-yA large language model-based approach to quantifying the effects of social determinants in liver transplant decisions
A large language model-based approach to quantifying the effects of social determinants in liver transplant decisions
npj Digital Medicine, Published online: 17 November 2025; doi:10.1038/s41746-025-02025-y
A large language model-based approach to quantifying the effects of social determinants in liver transplant decisions-
Nature - Issue - nature.com science feeds
-
The future of AI
Nature, Published online: 14 November 2025; doi:10.1038/d41586-025-03701-5Artificial intelligence is flying high. Nature asked leading innovators what they think will happen next.
The future of AI
Nature, Published online: 14 November 2025; doi:10.1038/d41586-025-03701-5
Artificial intelligence is flying high. Nature asked leading innovators what they think will happen next.-
Pulmonary nodule
-
Leveraging Artificial Intelligence to Transform Thoracic Radiology for Lung Nodules and Lung Cancer: Applications, Challenges, and Future Directions
J Thorac Imaging. 2026 Mar 1;41(2):e0866. doi: 10.1097/RTI.0000000000000866.ABSTRACTThis review traces the historical path of artificial intelligence (AI) methods that have been applied to medical image interpretation. Early AI approaches, which were based on clinical expertise and domain-specific medical knowledge, established the basis for data-driven methods, initiating the radiomics era and leading to the widespread use of deep learning in medical imaging. More recently, transformer architec
Leveraging Artificial Intelligence to Transform Thoracic Radiology for Lung Nodules and Lung Cancer: Applications, Challenges, and Future Directions
J Thorac Imaging. 2026 Mar 1;41(2):e0866. doi: 10.1097/RTI.0000000000000866.
ABSTRACT
This review traces the historical path of artificial intelligence (AI) methods that have been applied to medical image interpretation. Early AI approaches, which were based on clinical expertise and domain-specific medical knowledge, established the basis for data-driven methods, initiating the radiomics era and leading to the widespread use of deep learning in medical imaging. More recently, transformer architectures-originally developed for natural language processing-have been adapted for medical image analysis. In the first section, we explore the literature on the use of AI, specifically addressing lung nodules and lung cancer. AI has been effective in detecting lung nodules, evaluating their characteristics, and predicting cancer risk, while also addressing technical issues like kernel conversion. In lung cancer, AI has been applied to various clinical needs, including prognosis evaluation, mutation identification, treatment response analysis, operability prediction, treatment-related pneumonitis, and clinical information extraction. In the following section, we explore foundation models, multimodal AI, and a multiomic approach in the field of lung nodules and lung cancer. Finally, as AI models continue to evolve, so too must the approaches for evaluating their real-world utility; thus, we outline relevant methods for evaluating the performance and application of AI in thoracic radiology.
PMID:41246950 | DOI:10.1097/RTI.0000000000000866