Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
Semantic Laundering in AI Agent Architectures: Why Tool Boundaries Do Not Confer Epistemic Warrant
arXiv:2601.08333v1 Announce Type: new Abstract: LLM-based agent architectures systematically conflate information transport mechanisms with epistemic justification mechanisms. We formalize this class of architectural failures as semantic laundering: a pattern where propositions with absent or weak warrant are accepted by the system as admissible by crossing architecturally trusted interfaces. We show that semantic laundering constitutes an architectural realization of the Gettier problem: propo
-
cs.AI, q-bio.NC updates on arXiv.org
-
WaterCopilot: An AI-Driven Virtual Assistant for Water Management
arXiv:2601.08559v1 Announce Type: new Abstract: Sustainable water resource management in transboundary river basins is challenged by fragmented data, limited real-time access, and the complexity of integrating diverse information sources. This paper presents WaterCopilot-an AI-driven virtual assistant developed through collaboration between the International Water Management Institute (IWMI) and Microsoft Research for the Limpopo River Basin (LRB) to bridge these gaps through a unified, interac
WaterCopilot: An AI-Driven Virtual Assistant for Water Management
-
cs.AI, q-bio.NC updates on arXiv.org
-
GI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards
arXiv:2601.08183v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) show promise in gastroenterology, yet their performance against comprehensive clinical workflows and human benchmarks remains unverified. To systematically evaluate state-of-the-art MLLMs across a panoramic gastrointestinal endoscopy workflow and determine their clinical utility compared with human endoscopists. We constructed GI-Bench, a benchmark encompassing 20 fine-grained lesion categories. Twelve ML
GI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards
-
cs.AI, q-bio.NC updates on arXiv.org
-
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
arXiv:2601.08634v1 Announce Type: cross Abstract: While recent research has systematically documented political orientation in large language models (LLMs), existing evaluations rely primarily on direct probing or demographic persona engineering to surface ideological biases. In social psychology, however, political ideology is also understood as a downstream consequence of fundamental moral intuitions. In this work, we investigate the causal relationship between moral values and political posi
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
-
cs.AI, q-bio.NC updates on arXiv.org
-
ISLA: A U-Net for MRI-based acute ischemic stroke lesion segmentation with deep supervision, attention, domain adaptation, and ensemble learning
arXiv:2601.08732v1 Announce Type: cross Abstract: Accurate delineation of acute ischemic stroke lesions in MRI is a key component of stroke diagnosis and management. In recent years, deep learning models have been successfully applied to the automatic segmentation of such lesions. While most proposed architectures are based on the U-Net framework, they primarily differ in their choice of loss functions and in the use of deep supervision, residual connections, and attention mechanisms. Moreover,
ISLA: A U-Net for MRI-based acute ischemic stroke lesion segmentation with deep supervision, attention, domain adaptation, and ensemble learning
-
cs.AI, q-bio.NC updates on arXiv.org
-
MCP Bridge: A Lightweight, LLM-Agnostic RESTful Proxy for Model Context Protocol Servers
arXiv:2504.08999v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly augmented with external tools through standardized interfaces like the Model Context Protocol (MCP). However, current MCP implementations face critical limitations: they typically require local process execution through STDIO transports, making them impractical for resource-constrained environments like mobile devices, web browsers, and edge computing. We present MCP Bridge, a lightweight RES
MCP Bridge: A Lightweight, LLM-Agnostic RESTful Proxy for Model Context Protocol Servers
-
cs.AI, q-bio.NC updates on arXiv.org
-
Efficient and Reproducible Biomedical Question Answering using Retrieval Augmented Generation
arXiv:2505.07917v2 Announce Type: replace-cross Abstract: Biomedical question-answering (QA) systems require effective retrieval and generation components to ensure accuracy, efficiency, and scalability. This study systematically examines a Retrieval-Augmented Generation (RAG) system for biomedical QA, evaluating retrieval strategies and response time trade-offs. We first assess state-of-the-art retrieval methods, including BM25, BioBERT, MedCPT, and a hybrid approach, alongside common data sto
Efficient and Reproducible Biomedical Question Answering using Retrieval Augmented Generation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Aligning Trustworthy AI with Democracy: A Dual Taxonomy of Opportunities and Risks
arXiv:2505.13565v2 Announce Type: replace-cross Abstract: Artificial Intelligence (AI) poses both significant risks and valuable opportunities for democratic governance. This paper introduces a dual taxonomy to evaluate AI's complex relationship with democracy: the AI Risks to Democracy (AIRD) taxonomy, which identifies how AI can undermine core democratic principles such as autonomy, fairness, and trust; and the AI's Positive Contributions to Democracy (AIPD) taxonomy, which highlights AI's po
Aligning Trustworthy AI with Democracy: A Dual Taxonomy of Opportunities and Risks
-
STAT

-
STAT+: On Day 2 of JPM, Gilead lays outs it next test, a VC looks to raise funds, and one firm has FDA whiplash
This is the online version of The Readout, STAT’s flagship biotech newsletter. Sign up to get it in your inbox. You’re back. We’re sort of back. It’s Day 2 of JPM and we’re definitely not exhausted or delirious yet.This is Elaine Chen, Adam Feuerstein, Matt Herper, and Allison DeAngelis again. We’ve got a lot more news today, so let’s get to it. The next test for Kite Pharma — and Gilead It’s anito-cel, the CAR-T therapy for multiple myeloma that Gilead is developing in partnership wi
STAT+: On Day 2 of JPM, Gilead lays outs it next test, a VC looks to raise funds, and one firm has FDA whiplash
This is the online version of The Readout, STAT’s flagship biotech newsletter. Sign up to get it in your inbox.
You’re back. We’re sort of back. It’s Day 2 of JPM and we’re definitely not exhausted or delirious yet.
This is Elaine Chen, Adam Feuerstein, Matt Herper, and Allison DeAngelis again. We’ve got a lot more news today, so let’s get to it.
The next test for Kite Pharma — and Gilead
It’s anito-cel, the CAR-T therapy for multiple myeloma that Gilead is developing in partnership with Arcellx. Gilead submitted the therapy to the FDA sometime before the end of December, Cindy Perettie, executive vice president of Kite Pharma, the cell therapy division of Gilead, told STAT at a Gilead media breakfast.
Continue to STAT+ to read the full story…


© Alex Hogan/STAT
-
TechCrunch
-
Doctors think AI has a place in healthcare — but maybe not as a chatbot
OpenAI and Anthropic have each launched healthcare-focused products over the last week.
Doctors think AI has a place in healthcare — but maybe not as a chatbot
-
STAT

-
STAT+: OpenEvidence promises ‘medical super-intelligence’ at JPM. What is it?
You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday. OpenEvidence promises ‘medical super-intelligence’ It will be hard for OpenEvidence to top its 2025. The company announced nearly $500 million in funding last year and seemingly overnight became a go-to tool in the medical profession. A slide during the company’s Monday JPM presentation claims t
STAT+: OpenEvidence promises ‘medical super-intelligence’ at JPM. What is it?
You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday.
OpenEvidence promises ‘medical super-intelligence’
It will be hard for OpenEvidence to top its 2025. The company announced nearly $500 million in funding last year and seemingly overnight became a go-to tool in the medical profession. A slide during the company’s Monday JPM presentation claims that queries to the company’s clinical evidence chatbot grew from 2.6 million in 2024 to 17.9 million in December 2025, with well over 100 million queries for the year.
The company also revealed it will be launching “medical super-intelligence.” What does that mean? Katie Palmer explains in a new story.
Continue to STAT+ to read the full story…


© Adobe
-
cs.AI, q-bio.NC updates on arXiv.org
-
AI Safeguards, Generative AI and the Pandora Box: AI Safety Measures to Protect Businesses and Personal Reputation
arXiv:2601.06197v1 Announce Type: new Abstract: Generative AI has unleashed the power of content generation and it has also unwittingly opened the pandora box of realistic deepfake causing a number of social hazards and harm to businesses and personal reputation. The investigation & ramification of Generative AI technology across industries, the resolution & hybridization detection techniques using neural networks allows flagging of the content. Good detection techniques & flagging
AI Safeguards, Generative AI and the Pandora Box: AI Safety Measures to Protect Businesses and Personal Reputation
-
cs.AI, q-bio.NC updates on arXiv.org
-
ConSensus: Multi-Agent Collaboration for Multimodal Sensing
arXiv:2601.06453v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly grounded in sensor data to perceive and reason about human physiology and the physical world. However, accurately interpreting heterogeneous multimodal sensor data remains a fundamental challenge. We show that a single monolithic LLM often fails to reason coherently across modalities, leading to incomplete interpretations and prior-knowledge bias. We introduce ConSensus, a training-free multi-agent col
ConSensus: Multi-Agent Collaboration for Multimodal Sensing
-
cs.AI, q-bio.NC updates on arXiv.org
-
Why Slop Matters
arXiv:2601.06060v1 Announce Type: cross Abstract: AI-generated "slop" is often seen as digital pollution. We argue that this dismissal of the topic risks missing important aspects of AI Slop that deserve rigorous study. AI Slop serves a social function: it offers a supply-side solution to a variety of problems in cultural and economic demand - that, collectively, people want more content than humans can supply. We also argue that AI Slop is not mere digital detritus but has its own aesthetic va
Why Slop Matters
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Patient/Industry Trade-off in Medical Artificial Intelligence
arXiv:2601.06144v1 Announce Type: cross Abstract: Artificial intelligence (AI) in healthcare has led to many promising developments; however, increasingly, AI research is funded by the private sector leading to potential trade-offs between benefits to patients and benefits to industry. Health AI practitioners should prioritize successful adaptation into clinical practice in order to provide meaningful benefits to patients, but translation usually requires collaboration with industry. We discuss
The Patient/Industry Trade-off in Medical Artificial Intelligence
-
cs.AI, q-bio.NC updates on arXiv.org
-
$\texttt{AMEND++}$: Benchmarking Eligibility Criteria Amendments in Clinical Trials
arXiv:2601.06300v1 Announce Type: cross Abstract: Clinical trial amendments frequently introduce delays, increased costs, and administrative burden, with eligibility criteria being the most commonly amended component. We introduce \textit{eligibility criteria amendment prediction}, a novel NLP task that aims to forecast whether the eligibility criteria of an initial trial protocol will undergo future amendments. To support this task, we release $\texttt{AMEND++}$, a benchmark suite comprising t
$\texttt{AMEND++}$: Benchmarking Eligibility Criteria Amendments in Clinical Trials
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Large-Scale Study on the Development and Issues of Multi-Agent AI Systems
arXiv:2601.07136v1 Announce Type: cross Abstract: The rapid emergence of multi-agent AI systems (MAS), including LangChain, CrewAI, and AutoGen, has shaped how large language model (LLM) applications are developed and orchestrated. However, little is known about how these systems evolve and are maintained in practice. This paper presents the first large-scale empirical study of open-source MAS, analyzing over 42K unique commits and over 4.7K resolved issues across eight leading systems. Our ana
A Large-Scale Study on the Development and Issues of Multi-Agent AI Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
Learning from Reasoning Failures via Synthetic Data Generation
arXiv:2504.14523v2 Announce Type: replace Abstract: Training models on synthetic data has emerged as an increasingly important strategy for improving the performance of generative AI. This approach is particularly helpful for large multimodal models (LMMs) due to the relative scarcity of high-quality paired image-text data compared to language-only data. While a variety of methods have been proposed for generating large multimodal datasets, they do not tailor the synthetic data to address speci
Learning from Reasoning Failures via Synthetic Data Generation
-
cs.AI, q-bio.NC updates on arXiv.org
-
FairMedQA: Benchmarking Bias in Large Language Models for Medical Question Answering
arXiv:2505.19562v2 Announce Type: replace Abstract: Large language models (LLMs) are approaching expert-level performance in medical question answering (QA), demonstrating strong potential to improve public healthcare. However, underlying biases related to sensitive attributes such as sex and race pose life-critical risks. The extent to which such sensitive attributes affect diagnosis remains an open question and requires comprehensive empirical investigation. Additionally, even the latest Coun
FairMedQA: Benchmarking Bias in Large Language Models for Medical Question Answering
-
cs.AI, q-bio.NC updates on arXiv.org
-
From Wearables to Warnings: Predicting Pain Spikes in Patients with Opioid Use Disorder
arXiv:2511.19577v2 Announce Type: replace Abstract: Chronic pain (CP) and opioid use disorder (OUD) are common and interrelated chronic medical conditions. Currently, there is a paucity of evidence-based integrated treatments for CP and OUD among individuals receiving medication for opioid use disorder (MOUD). Wearable devices have the potential to monitor complex patient information and inform treatment development for persons with OUD and CP, including pain variability (e.g., exacerbations of