Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
Small Language Models for Efficient Agentic Tool Calling: Outperforming Large Models with Targeted Fine-tuning
arXiv:2512.15943v1 Announce Type: new Abstract: As organizations scale adoption of generative AI, model cost optimization and operational efficiency have emerged as critical factors determining sustainability and accessibility. While Large Language Models (LLMs) demonstrate impressive capabilities across diverse tasks, their extensive computational requirements make them cost-prohibitive for routine enterprise use. This limitation motivates the exploration of Small Language Models (SLMs), which
-
cs.AI, q-bio.NC updates on arXiv.org
-
Distributional AGI Safety
arXiv:2512.16856v1 Announce Type: new Abstract: AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an eventual emergence of a monolithic Artificial General Intelligence (AGI). The alternative AGI emergence hypothesis, where general capability levels are first manifested through coordination in groups of sub-AGI individual agents with complementary skills and affordances, has received far less attention.
Distributional AGI Safety
-
cs.AI, q-bio.NC updates on arXiv.org
-
Emergent Bias and Fairness in Multi-Agent Decision Systems
arXiv:2512.16433v1 Announce Type: cross Abstract: Multi-agent systems have demonstrated the ability to improve performance on a variety of predictive tasks by leveraging collaborative decision making. However, the lack of effective evaluation methodologies has made it difficult to estimate the risk of bias, making deployment of such systems unsafe in high stakes domains such as consumer finance, where biased decisions can translate directly into regulatory breaches and financial loss. To addres
Emergent Bias and Fairness in Multi-Agent Decision Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
AI4EOSC: a Federated Cloud Platform for Artificial Intelligence in Scientific Research
arXiv:2512.16455v1 Announce Type: cross Abstract: In this paper, we describe a federated compute platform dedicated to support Artificial Intelligence in scientific workloads. Putting the effort into reproducible deployments, it delivers consistent, transparent access to a federation of physically distributed e-Infrastructures. Through a comprehensive service catalogue, the platform is able to offer an integrated user experience covering the full Machine Learning lifecycle, including model deve
AI4EOSC: a Federated Cloud Platform for Artificial Intelligence in Scientific Research
-
cs.AI, q-bio.NC updates on arXiv.org
-
From Facts to Conclusions : Integrating Deductive Reasoning in Retrieval-Augmented LLMs
arXiv:2512.16795v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) grounds large language models (LLMs) in external evidence, but fails when retrieved sources conflict or contain outdated or subjective information. Prior work address these issues independently but lack unified reasoning supervision. We propose a reasoning-trace-augmented RAG framework that adds structured, interpretable reasoning across three stages : (1) document-level adjudication, (2) conflict analysis, a
From Facts to Conclusions : Integrating Deductive Reasoning in Retrieval-Augmented LLMs
-
cs.AI, q-bio.NC updates on arXiv.org
-
Constitutional Law and AI Governance: Constraints on Model Licensing and Research Classification
arXiv:2509.05361v2 Announce Type: replace-cross Abstract: Transformative AI systems may pose unprecedented catastrophic risks, but the U.S. Constitution places significant constraints on the government's ability to govern this technology. This paper examines how the First Amendment, administrative law, and the Fourteenth Amendment shape the legal vulnerability of two regulatory proposals: model licensing and AI research classification. While the First Amendment may provide some degree of protec
Constitutional Law and AI Governance: Constraints on Model Licensing and Research Classification
-
cs.AI, q-bio.NC updates on arXiv.org
-
First, do NOHARM: towards clinically safe large language models
arXiv:2512.01241v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized. We present NOHARM (Numerous Options Harm Assessment for Risk in Medicine), a benchmark using 100 real primary care-to-specialist consultation cases to measure frequency and severity of harm from LLM-generated medical recommendations. NOHARM covers 10 specialties, with 12,747 expert
First, do NOHARM: towards clinically safe large language models
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Decision-Theoretic Approach for Managing Misalignment
arXiv:2512.15584v1 Announce Type: new Abstract: When should we delegate decisions to AI systems? While the value alignment literature has developed techniques for shaping AI values, less attention has been paid to how to determine, under uncertainty, when imperfect alignment is good enough to justify delegation. We argue that rational delegation requires balancing an agent's value (mis)alignment with its epistemic accuracy and its reach (the acts it has available). This paper introduces a forma
A Decision-Theoretic Approach for Managing Misalignment
-
cs.AI, q-bio.NC updates on arXiv.org
-
DrugRAG: Enhancing Pharmacy LLM Performance Through A Novel Retrieval-Augmented Generation Pipeline
arXiv:2512.14896v1 Announce Type: cross Abstract: Objectives: To evaluate large language model (LLM) performance on pharmacy licensure-style question-answering (QA) tasks and develop an external knowledge integration method to improve their accuracy. Methods: We benchmarked eleven existing LLMs with varying parameter sizes (8 billion to 70+ billion) using a 141-question pharmacy dataset. We measured baseline accuracy for each model without modification. We then developed a three-step retrieva
DrugRAG: Enhancing Pharmacy LLM Performance Through A Novel Retrieval-Augmented Generation Pipeline
-
cs.AI, q-bio.NC updates on arXiv.org
-
MedChat: A Multi-Agent Framework for Multimodal Diagnosis with Large Language Models
arXiv:2506.07400v3 Announce Type: replace-cross Abstract: The integration of deep learning-based glaucoma detection with large language models (LLMs) presents an automated strategy to mitigate ophthalmologist shortages and improve clinical reporting efficiency. However, applying general LLMs to medical imaging remains challenging due to hallucinations, limited interpretability, and insufficient domain-specific medical knowledge, which can potentially reduce clinical accuracy. Although recent ap
MedChat: A Multi-Agent Framework for Multimodal Diagnosis with Large Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Multimodal Foundation Models for Early Disease Detection
arXiv:2510.01899v2 Announce Type: replace-cross Abstract: Healthcare data now span EHRs, medical imaging, genomics, and wearable sensors, but most diagnostic models still process these modalities in isolation. This limits their ability to capture early, cross-modal disease signatures. This paper introduces a multimodal foundation model built on a transformer architecture that integrates heterogeneous clinical data through modality-specific encoders and cross-modal attention. Each modality is ma
Multimodal Foundation Models for Early Disease Detection
-
(Multiomics OR Omics) AND (Lung OR gastric OR Hepatocellular)
-
A novel statistical feature selection framework for biomarker discovery and cancer classification via multiomics integration
BMC Med Res Methodol. 2025 Dec 17. doi: 10.1186/s12874-025-02713-z. Online ahead of print.ABSTRACTBACKGROUND: Early cancer diagnosis is essential for improving prognosis and guiding treatment. However, the high dimensionality and complexity of omics data present major challenges. Computational approaches that extract stable biomarkers and enable reliable classification across cancer types and stages are needed.METHODS: A novel feature selection method, sDCFE (synergistic Discriminative Cluster-b
A novel statistical feature selection framework for biomarker discovery and cancer classification via multiomics integration
BMC Med Res Methodol. 2025 Dec 17. doi: 10.1186/s12874-025-02713-z. Online ahead of print.
ABSTRACT
BACKGROUND: Early cancer diagnosis is essential for improving prognosis and guiding treatment. However, the high dimensionality and complexity of omics data present major challenges. Computational approaches that extract stable biomarkers and enable reliable classification across cancer types and stages are needed.
METHODS: A novel feature selection method, sDCFE (synergistic Discriminative Cluster-based Feature Extraction), was developed by extending Fisher-like variance analysis with a median absolute deviation (MAD) regularization term and a cluster separation component to enhance robustness and interpretability. Features selected by sDCFE were compared with those obtained from XGBoost, and the intersected set of 82 genes was evaluated through functional enrichment (KEGG, Reactome, GO BP), survival analysis (Kaplan-Meier, Cox regression), and biomarker novelty assessment against six external resources. Hybrid classification models integrating XGBoost, sDCFE, and deep learning were applied to pancancer classification, and the framework was further extended to lung squamous cell carcinoma (LUSC) staging using RNA-seq and methylation data.
RESULTS: The overlap between sDCFE and XGBoost yielded 82 candidate biomarkers enriched in cancer-related pathways, including cell cycle regulation, immune signalling, and DNA repair. Novelty assessment stratified these genes into established, emerging, and novel categories. Six genes-HFE2, LOC339674, SERINC2, SFTA3, SOX2OT, and ACPP-emerged as the most promising candidates, supported by enrichment and survival associations across multiple cancers. The hybrid model achieved near-perfect pancancer classification on TCGA (accuracy = 99.3%, MCC = 0.992, AUC = 1.0) and demonstrated strong generalizability on PCAWG (accuracy = 94%, MCC = 0.929, AUC = 0.997). In the LUSC staging task, multiomics integration improved classification performance: the CNN-based model reached 84% accuracy, while logistic regression applied to sDCFE-ranked features achieved 88.5% accuracy with superior calibration, highlighting the robustness of the selected features.
CONCLUSION: sDCFE provides a principled extension of Fisher-like methods, enabling stable and interpretable biomarker selection. When combined with XGBoost and deep learning, the framework achieves highly accurate and biologically grounded cancer classification across both cancer types and stages. The identification of novel and prognostic biomarkers, including HFE2, LOC339674, SERINC2, SFTA3, SOX2OT, and ACPP, underscores its translational potential. These results position the framework as a promising precision oncology tool to support early diagnosis, risk stratification, and treatment decision-making.
PMID:41408184 | DOI:10.1186/s12874-025-02713-z
-
cs.AI, q-bio.NC updates on arXiv.org
-
Enhancing Transparency and Traceability in Healthcare AI: The AI Product Passport
arXiv:2512.13702v1 Announce Type: cross Abstract: Objective: To develop the AI Product Passport, a standards-based framework improving transparency, traceability, and compliance in healthcare AI via lifecycle-based documentation. Materials and Methods: The AI Product Passport was developed within the AI4HF project, focusing on heart failure AI tools. We analyzed regulatory frameworks (EU AI Act, FDA guidelines) and existing standards to design a relational data model capturing metadata across A
Enhancing Transparency and Traceability in Healthcare AI: The AI Product Passport
-
cs.AI, q-bio.NC updates on arXiv.org
-
Graph AI generates neurological hypotheses validated in molecular, organoid, and clinical systems
arXiv:2512.13724v1 Announce Type: cross Abstract: Neurological diseases are the leading global cause of disability, yet most lack disease-modifying treatments. We present PROTON, a heterogeneous graph transformer that generates testable hypotheses across molecular, organoid, and clinical systems. To evaluate PROTON, we apply it to Parkinson's disease (PD), bipolar disorder (BD), and Alzheimer's disease (AD). In PD, PROTON linked genetic risk loci to genes essential for dopaminergic neuron survi
Graph AI generates neurological hypotheses validated in molecular, organoid, and clinical systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
Criminal Liability in AI-Enabled Autonomous Vehicles: A Comparative Study
arXiv:2512.14330v1 Announce Type: cross Abstract: AI revolutionizes transportation through autonomous vehicles (AVs) but introduces complex criminal liability issues regarding infractions. This study employs a comparative legal analysis of primary statutes, real-world liability claims, and academic literature across the US, Germany, UK, China, and India; jurisdictions selected for their technological advancement and contrasting regulatory approaches. The research examines the attribution of hum
Criminal Liability in AI-Enabled Autonomous Vehicles: A Comparative Study
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Multicenter Benchmark of Multiple Instance Learning Models for Lymphoma Subtyping from HE-stained Whole Slide Images
arXiv:2512.14640v1 Announce Type: cross Abstract: Timely and accurate lymphoma diagnosis is essential for guiding cancer treatment. Standard diagnostic practice combines hematoxylin and eosin (HE)-stained whole slide images with immunohistochemistry, flow cytometry, and molecular genetic tests to determine lymphoma subtypes, a process requiring costly equipment, skilled personnel, and causing treatment delays. Deep learning methods could assist pathologists by extracting diagnostic information
A Multicenter Benchmark of Multiple Instance Learning Models for Lymphoma Subtyping from HE-stained Whole Slide Images
-
cs.AI, q-bio.NC updates on arXiv.org
-
COMMA: A Communicative Multimodal Multi-Agent Benchmark
arXiv:2410.07553v5 Announce Type: replace Abstract: The rapid advances of multimodal agents built on large foundation models have largely overlooked their potential for language-based communication between agents in collaborative tasks. This oversight presents a critical gap in understanding their effectiveness in real-world deployments, particularly when communicating with humans. Existing agentic benchmarks fail to address key aspects of inter-agent communication and collaboration, particular
COMMA: A Communicative Multimodal Multi-Agent Benchmark
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Knowledge Graph-based Retrieval-Augmented Generation Framework for Algorithm Selection in the Facility Layout Problem
arXiv:2509.18054v2 Announce Type: replace-cross Abstract: Selecting a solution algorithm for the Facility Layout Problem (FLP), an NP-hard optimization problem with multiobjective trade-off, is a complex task that requires deep expert knowledge. The performance of a given algorithm depends on the specific characteristics of the problem, such as the number of facilities, objectives, and constraints. This creates a need for a data-driven recommendation method to guide algorithm selection in autom
A Knowledge Graph-based Retrieval-Augmented Generation Framework for Algorithm Selection in the Facility Layout Problem
-
cs.AI, q-bio.NC updates on arXiv.org
-
Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems
arXiv:2512.12791v2 Announce Type: replace-cross Abstract: Recent advances in agentic AI have shifted the focus from standalone Large Language Models (LLMs) to integrated systems that combine LLMs with tools, memory, and other agents to perform complex tasks. These multi-agent architectures enable coordinated reasoning, planning, and execution across diverse domains, allowing agents to collaboratively automate complex workflows. Despite these advances, evaluation and assessment of LLM agents and
Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems
-
Nature - Issue - nature.com science feeds
-
Immunological sin: how a person’s earliest flu infections dictate life-long immunity
Nature, Published online: 17 December 2025; doi:10.1038/d41586-025-03606-3Researchers are striving to understand the impact a phenomenon known as original antigenic sin has on immunity to the virus.
Immunological sin: how a person’s earliest flu infections dictate life-long immunity
Nature, Published online: 17 December 2025; doi:10.1038/d41586-025-03606-3
Researchers are striving to understand the impact a phenomenon known as original antigenic sin has on immunity to the virus.