Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
Semantic Laundering in AI Agent Architectures: Why Tool Boundaries Do Not Confer Epistemic Warrant
arXiv:2601.08333v1 Announce Type: new Abstract: LLM-based agent architectures systematically conflate information transport mechanisms with epistemic justification mechanisms. We formalize this class of architectural failures as semantic laundering: a pattern where propositions with absent or weak warrant are accepted by the system as admissible by crossing architecturally trusted interfaces. We show that semantic laundering constitutes an architectural realization of the Gettier problem: propo
-
cs.AI, q-bio.NC updates on arXiv.org
-
WaterCopilot: An AI-Driven Virtual Assistant for Water Management
arXiv:2601.08559v1 Announce Type: new Abstract: Sustainable water resource management in transboundary river basins is challenged by fragmented data, limited real-time access, and the complexity of integrating diverse information sources. This paper presents WaterCopilot-an AI-driven virtual assistant developed through collaboration between the International Water Management Institute (IWMI) and Microsoft Research for the Limpopo River Basin (LRB) to bridge these gaps through a unified, interac
WaterCopilot: An AI-Driven Virtual Assistant for Water Management
-
cs.AI, q-bio.NC updates on arXiv.org
-
All Required, In Order: Phase-Level Evaluation for AI-Human Dialogue in Healthcare and Beyond
arXiv:2601.08690v1 Announce Type: new Abstract: Conversational AI is starting to support real clinical work, but most evaluation methods miss how compliance depends on the full course of a conversation. We introduce Obligatory-Information Phase Structured Compliance Evaluation (OIP-SCE), an evaluation method that checks whether every required clinical obligation is met, in the right order, with clear evidence for clinicians to review. This makes complex rules practical and auditable, helping cl
All Required, In Order: Phase-Level Evaluation for AI-Human Dialogue in Healthcare and Beyond
-
cs.AI, q-bio.NC updates on arXiv.org
-
GI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards
arXiv:2601.08183v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) show promise in gastroenterology, yet their performance against comprehensive clinical workflows and human benchmarks remains unverified. To systematically evaluate state-of-the-art MLLMs across a panoramic gastrointestinal endoscopy workflow and determine their clinical utility compared with human endoscopists. We constructed GI-Bench, a benchmark encompassing 20 fine-grained lesion categories. Twelve ML
GI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards
-
cs.AI, q-bio.NC updates on arXiv.org
-
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
arXiv:2601.08634v1 Announce Type: cross Abstract: While recent research has systematically documented political orientation in large language models (LLMs), existing evaluations rely primarily on direct probing or demographic persona engineering to surface ideological biases. In social psychology, however, political ideology is also understood as a downstream consequence of fundamental moral intuitions. In this work, we investigate the causal relationship between moral values and political posi
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
-
cs.AI, q-bio.NC updates on arXiv.org
-
ISLA: A U-Net for MRI-based acute ischemic stroke lesion segmentation with deep supervision, attention, domain adaptation, and ensemble learning
arXiv:2601.08732v1 Announce Type: cross Abstract: Accurate delineation of acute ischemic stroke lesions in MRI is a key component of stroke diagnosis and management. In recent years, deep learning models have been successfully applied to the automatic segmentation of such lesions. While most proposed architectures are based on the U-Net framework, they primarily differ in their choice of loss functions and in the use of deep supervision, residual connections, and attention mechanisms. Moreover,
ISLA: A U-Net for MRI-based acute ischemic stroke lesion segmentation with deep supervision, attention, domain adaptation, and ensemble learning
-
cs.AI, q-bio.NC updates on arXiv.org
-
DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extraction
arXiv:2509.14507v2 Announce Type: replace Abstract: Natural Language to SQL (NL2SQL) provides a new model-centric paradigm that simplifies database access for non-technical users by converting natural language queries into SQL commands. Recent advancements, particularly those integrating Retrieval-Augmented Generation (RAG) and Chain-of-Thought (CoT) reasoning, have made significant strides in enhancing NL2SQL performance. However, challenges such as inaccurate task decomposition and keyword ex
DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extraction
-
cs.AI, q-bio.NC updates on arXiv.org
-
Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
arXiv:2512.20387v4 Announce Type: replace Abstract: We propose a Vision-Language Simulation Model (VLSM) that unifies visual and textual understanding to synthesize executable FlexScript from layout sketches and natural-language prompts, enabling cross-modal reasoning for industrial simulation systems. To support this new paradigm, the study constructs the first large-scale dataset for generative digital twins, comprising over 120,000 prompt-sketch-code triplets that enable multimodal learning
Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
Explainable Molecular Property Prediction: Aligning Chemical Concepts with Predictions via Language Models
arXiv:2405.16041v4 Announce Type: replace-cross Abstract: Providing explainable molecular property predictions is critical for many scientific domains, such as drug discovery and material science. Though transformer-based language models have shown great potential in accurate molecular property prediction, they neither provide chemically meaningful explanations nor faithfully reveal the molecular structure-property relationships. In this work, we develop a framework for explainable molecular pr
Explainable Molecular Property Prediction: Aligning Chemical Concepts with Predictions via Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Hallucination, reliability, and the role of generative AI in science
arXiv:2504.08526v2 Announce Type: replace-cross Abstract: Generative AI increasingly supports scientific inference, from protein structure prediction to weather forecasting. Yet its distinctive failure mode, hallucination, raises epistemic alarm bells. I argue that this failure mode can be addressed by shifting from data-centric to phenomenon-centric assessment. Through case studies of AlphaFold and GenCast, I show how scientific workflows discipline generative models through theory-guided trai
Hallucination, reliability, and the role of generative AI in science
-
cs.AI, q-bio.NC updates on arXiv.org
-
Efficient and Reproducible Biomedical Question Answering using Retrieval Augmented Generation
arXiv:2505.07917v2 Announce Type: replace-cross Abstract: Biomedical question-answering (QA) systems require effective retrieval and generation components to ensure accuracy, efficiency, and scalability. This study systematically examines a Retrieval-Augmented Generation (RAG) system for biomedical QA, evaluating retrieval strategies and response time trade-offs. We first assess state-of-the-art retrieval methods, including BM25, BioBERT, MedCPT, and a hybrid approach, alongside common data sto
Efficient and Reproducible Biomedical Question Answering using Retrieval Augmented Generation
-
STAT

-
STAT+: On Day 2 of JPM, Gilead lays outs it next test, a VC looks to raise funds, and one firm has FDA whiplash
This is the online version of The Readout, STAT’s flagship biotech newsletter. Sign up to get it in your inbox. You’re back. We’re sort of back. It’s Day 2 of JPM and we’re definitely not exhausted or delirious yet.This is Elaine Chen, Adam Feuerstein, Matt Herper, and Allison DeAngelis again. We’ve got a lot more news today, so let’s get to it. The next test for Kite Pharma — and Gilead It’s anito-cel, the CAR-T therapy for multiple myeloma that Gilead is developing in partnership wi
STAT+: On Day 2 of JPM, Gilead lays outs it next test, a VC looks to raise funds, and one firm has FDA whiplash
This is the online version of The Readout, STAT’s flagship biotech newsletter. Sign up to get it in your inbox.
You’re back. We’re sort of back. It’s Day 2 of JPM and we’re definitely not exhausted or delirious yet.
This is Elaine Chen, Adam Feuerstein, Matt Herper, and Allison DeAngelis again. We’ve got a lot more news today, so let’s get to it.
The next test for Kite Pharma — and Gilead
It’s anito-cel, the CAR-T therapy for multiple myeloma that Gilead is developing in partnership with Arcellx. Gilead submitted the therapy to the FDA sometime before the end of December, Cindy Perettie, executive vice president of Kite Pharma, the cell therapy division of Gilead, told STAT at a Gilead media breakfast.
Continue to STAT+ to read the full story…


© Alex Hogan/STAT
-
npj Digital Medicine
-
Geometric multi-instance learning for weakly supervised gastric cancer segmentation
npj Digital Medicine, Published online: 13 January 2026; doi:10.1038/s41746-025-02287-6Geometric multi-instance learning for weakly supervised gastric cancer segmentation
Geometric multi-instance learning for weakly supervised gastric cancer segmentation
npj Digital Medicine, Published online: 13 January 2026; doi:10.1038/s41746-025-02287-6
Geometric multi-instance learning for weakly supervised gastric cancer segmentation-
cs.AI, q-bio.NC updates on arXiv.org
-
SafePro: Evaluating the Safety of Professional-Level AI Agents
arXiv:2601.06663v1 Announce Type: new Abstract: Large language model-based agents are rapidly evolving from simple conversational assistants into autonomous systems capable of performing complex, professional-level tasks in various domains. While these advancements promise significant productivity gains, they also introduce critical safety risks that remain under-explored. Existing safety evaluations primarily focus on simple, daily assistance tasks, failing to capture the intricate decision-ma
SafePro: Evaluating the Safety of Professional-Level AI Agents
-
cs.AI, q-bio.NC updates on arXiv.org
-
From "Thinking" to "Justifying": Aligning High-Stakes Explainability with Professional Communication Standards
arXiv:2601.07233v1 Announce Type: new Abstract: Explainable AI (XAI) in high-stakes domains should help stakeholders trust and verify system outputs. Yet Chain-of-Thought methods reason before concluding, and logical gaps or hallucinations can yield conclusions that do not reliably align with their rationale. Thus, we propose "Result -> Justify", which constrains the output communication to present a conclusion before its structured justification. We introduce SEF (Structured Explainability
From "Thinking" to "Justifying": Aligning High-Stakes Explainability with Professional Communication Standards
-
cs.AI, q-bio.NC updates on arXiv.org
-
Puzzle it Out: Local-to-Global World Model for Offline Multi-Agent Reinforcement Learning
arXiv:2601.07463v1 Announce Type: new Abstract: Offline multi-agent reinforcement learning (MARL) aims to solve cooperative decision-making problems in multi-agent systems using pre-collected datasets. Existing offline MARL methods primarily constrain training within the dataset distribution, resulting in overly conservative policies that struggle to generalize beyond the support of the data. While model-based approaches offer a promising solution by expanding the original dataset with syntheti
Puzzle it Out: Local-to-Global World Model for Offline Multi-Agent Reinforcement Learning
-
cs.AI, q-bio.NC updates on arXiv.org
-
From Augmentation to Symbiosis: A Review of Human-AI Collaboration Frameworks, Performance, and Perils
arXiv:2601.06030v1 Announce Type: cross Abstract: This paper offers a concise, 60-year synthesis of human-AI collaboration, from Licklider's ``man-computer symbiosis" (AI as colleague) and Engelbart's ``augmenting human intellect" (AI as tool) to contemporary poles: Human-Centered AI's ``supertool" and Symbiotic Intelligence's mutual-adaptation model. We formalize the mechanism for effective teaming as a causal chain: Explainable AI (XAI) -> co-adaptation -> shared mental models (SMMs). A
From Augmentation to Symbiosis: A Review of Human-AI Collaboration Frameworks, Performance, and Perils
-
cs.AI, q-bio.NC updates on arXiv.org
-
Why Slop Matters
arXiv:2601.06060v1 Announce Type: cross Abstract: AI-generated "slop" is often seen as digital pollution. We argue that this dismissal of the topic risks missing important aspects of AI Slop that deserve rigorous study. AI Slop serves a social function: it offers a supply-side solution to a variety of problems in cultural and economic demand - that, collectively, people want more content than humans can supply. We also argue that AI Slop is not mere digital detritus but has its own aesthetic va
Why Slop Matters
-
cs.AI, q-bio.NC updates on arXiv.org
-
Interoperability in AI Safety Governance: Ethics, Regulations, and Standards
arXiv:2601.06153v1 Announce Type: cross Abstract: This policy report draws on country studies from China, South Korea, Singapore, and the United Kingdom to identify effective tools and key barriers to interoperability in AI safety governance. It offers practical recommendations to support a globally informed yet locally grounded governance ecosystem. Interoperability is a central goal of AI governance, vital for reducing risks, fostering innovation, enhancing competitiveness, promoting standard
Interoperability in AI Safety Governance: Ethics, Regulations, and Standards
-
cs.AI, q-bio.NC updates on arXiv.org
-
Toward Safe and Responsible AI Agents: A Three-Pillar Model for Transparency, Accountability, and Trustworthiness
arXiv:2601.06223v1 Announce Type: cross Abstract: This paper presents a conceptual and operational framework for developing and operating safe and trustworthy AI agents based on a Three-Pillar Model grounded in transparency, accountability, and trustworthiness. Building on prior work in Human-in-the-Loop systems, reinforcement learning, and collaborative AI, the framework defines an evolutionary path toward autonomous agents that balances increasing automation with appropriate human oversight.