Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
Causal Graph Neural Networks for Healthcare
arXiv:2511.02531v1 Announce Type: cross Abstract: Healthcare artificial intelligence systems routinely fail when deployed across institutions, with documented performance drops and perpetuation of discriminatory patterns embedded in historical data. This brittleness stems, in part, from learning statistical associations rather than causal mechanisms. Causal graph neural networks address this triple crisis of distribution shift, discrimination, and inscrutability by combining graph-based represe
-
cs.AI, q-bio.NC updates on arXiv.org
-
TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models
arXiv:2511.02802v1 Announce Type: cross Abstract: Tabular foundation models represent a growing paradigm in structured data learning, extending the benefits of large-scale pretraining to tabular domains. However, their adoption remains limited due to heterogeneous preprocessing pipelines, fragmented APIs, inconsistent fine-tuning procedures, and the absence of standardized evaluation for deployment-oriented metrics such as calibration and fairness. We present TabTune, a unified library that sta
TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
How can we assess human-agent interactions? Case studies in software agent design
arXiv:2510.09801v2 Announce Type: replace Abstract: LLM-powered agents are both a promising new technology and a source of complexity, where choices about models, tools, and prompting can affect their usefulness. While numerous benchmarks measure agent accuracy across domains, they mostly assume full automation, failing to represent the collaborative nature of real-world use cases. In this paper, we make two major steps towards the rigorous assessment of human-agent interactions. First, we prop
How can we assess human-agent interactions? Case studies in software agent design
-
cs.AI, q-bio.NC updates on arXiv.org
-
AutoPDL: Automatic Prompt Optimization for LLM Agents
arXiv:2504.04365v5 Announce Type: replace-cross Abstract: The performance of large language models (LLMs) depends on how they are prompted, with choices spanning both the high-level prompting pattern (e.g., Zero-Shot, CoT, ReAct, ReWOO) and the specific prompt content (instructions and few-shot demonstrations). Manually tuning this combination is tedious, error-prone, and specific to a given LLM and task. Therefore, this paper proposes AutoPDL, an automated approach to discovering good LLM agen
AutoPDL: Automatic Prompt Optimization for LLM Agents
-
cs.AI, q-bio.NC updates on arXiv.org
-
Diffusion Models at the Drug Discovery Frontier: A Review on Generating Small Molecules versus Therapeutic Peptides
arXiv:2511.00209v1 Announce Type: cross Abstract: Diffusion models have emerged as a leading framework in generative modeling, showing significant potential to accelerate and transform the traditionally slow and costly process of drug discovery. This review provides a systematic comparison of their application in designing two principal therapeutic modalities: small molecules and therapeutic peptides. We analyze how a unified framework of iterative denoising is adapted to the distinct molecular
Diffusion Models at the Drug Discovery Frontier: A Review on Generating Small Molecules versus Therapeutic Peptides
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation
arXiv:2510.19755v3 Announce Type: replace-cross Abstract: Diffusion Models have become a cornerstone of modern generative AI for their exceptional generation quality and controllability. However, their inherent \textit{multi-step iterations} and \textit{complex backbone networks} lead to prohibitive computational overhead and generation latency, forming a major bottleneck for real-time applications. Although existing acceleration techniques have made progress, they still face challenges such as
A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation
-
Journal of Medical Internet Research
-
Generative Artificial Intelligence in Medical Education: Enhancing Critical Thinking or Undermining Cognitive Autonomy?
Generative artificial intelligence (GenAI) enables the production of coherent and contextually relevant text by processing large-scale linguistic datasets. Tools such as ChatGPT, Gemini, Claude, and LLaMA are increasingly integrated into medical education, assisting students with a range of tasks, including clinical reasoning, literature review, scientific writing, and formative assessment. Although these tools offer significant advantages in terms of productivity, personalization, and cognitive
Generative Artificial Intelligence in Medical Education: Enhancing Critical Thinking or Undermining Cognitive Autonomy?
-
(Multiomics OR Omics) AND (Pancreatic)
-
Artificial intelligence in pancreatitis: A narrative review on advancing precision diagnosis, prognosis, and therapeutic strategies
World J Gastroenterol. 2025 Oct 21;31(39):110971. doi: 10.3748/wjg.v31.i39.110971.ABSTRACTPancreatitis poses persistent diagnostic and therapeutic challenges due to its heterogeneous clinical presentation, variable disease course, and lack of targeted interventions. Conventional tools, such as serum enzymes, cross-sectional imaging and clinical scoring systems, often exhibit limited sensitivity and prognostic value, especially during early or atypical stages. Moreover, therapeutic development re
Artificial intelligence in pancreatitis: A narrative review on advancing precision diagnosis, prognosis, and therapeutic strategies
World J Gastroenterol. 2025 Oct 21;31(39):110971. doi: 10.3748/wjg.v31.i39.110971.
ABSTRACT
Pancreatitis poses persistent diagnostic and therapeutic challenges due to its heterogeneous clinical presentation, variable disease course, and lack of targeted interventions. Conventional tools, such as serum enzymes, cross-sectional imaging and clinical scoring systems, often exhibit limited sensitivity and prognostic value, especially during early or atypical stages. Moreover, therapeutic development remains slow, with limited progress toward personalized or mechanism-based strategies. These limitations highlight a critical need for integrative data-driven approaches. Artificial intelligence (AI) has emerged as a promising tool to enhance clinical decision-making in pancreatitis. This narrative review synthesizes recent progress in AI applications across three domains. First, AI-enabled diagnostic platforms incorporating radiomics, deep learning-based imaging analysis, and biomarker optimization have improved early detection and differentiation of pancreatic diseases. Second, AI-driven prognostic models now allow real-time severity prediction, complication forecasting, and recurrence risk assessment, some of which have been deployed in hospital information systems for intensive care units and mortality risk triage. Third, AI-assisted drug discovery and network pharmacology, particularly in combination with traditional Chinese medicine, have revealed novel therapeutic opportunities. Despite encouraging developments, challenges remain in data standardization, model transparency and clinical validation. A multidisciplinary strategy integrating omics data, longitudinal monitoring and pharmacological modeling may help bridge current gaps and advance precision medicine in pancreatitis care.
PMID:41180795 | PMC:PMC12576603 | DOI:10.3748/wjg.v31.i39.110971
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Denario project: Deep knowledge AI agents for scientific discovery
arXiv:2510.26887v1 Announce Type: new Abstract: We present Denario, an AI multi-agent system designed to serve as a scientific research assistant. Denario can perform many different tasks, such as generating ideas, checking the literature, developing research plans, writing and executing code, making plots, and drafting and reviewing a scientific paper. The system has a modular architecture, allowing it to handle specific tasks, such as generating an idea, or carrying out end-to-end scientific
The Denario project: Deep knowledge AI agents for scientific discovery
-
cs.AI, q-bio.NC updates on arXiv.org
-
Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement
arXiv:2510.27051v1 Announce Type: new Abstract: Enterprise AI agents must continuously adapt to maintain accuracy, reduce latency, and remain aligned with user needs. We present a practical implementation of a data flywheel in NVInfo AI, NVIDIA's Mixture-of-Experts (MoE) Knowledge Assistant serving over 30,000 employees. By operationalizing a MAPE-driven data flywheel, we built a closed-loop system that systematically addresses failures in retrieval-augmented generation (RAG) pipelines and enab
Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement
-
cs.AI, q-bio.NC updates on arXiv.org
-
Glia: A Human-Inspired AI for Automated Systems Design and Optimization
arXiv:2510.27176v1 Announce Type: new Abstract: Can an AI autonomously design mechanisms for computer systems on par with the creativity and reasoning of human experts? We present Glia, an AI architecture for networked systems design that uses large language models (LLMs) in a human-inspired, multi-agent workflow. Each agent specializes in reasoning, experimentation, and analysis, collaborating through an evaluation framework that grounds abstract reasoning in empirical feedback. Unlike prior M
Glia: A Human-Inspired AI for Automated Systems Design and Optimization
-
cs.AI, q-bio.NC updates on arXiv.org
-
An In-depth Study of LLM Contributions to the Bin Packing Problem
arXiv:2510.27353v1 Announce Type: new Abstract: Recent studies have suggested that Large Language Models (LLMs) could provide interesting ideas contributing to mathematical discovery. This claim was motivated by reports that LLM-based genetic algorithms produced heuristics offering new insights into the online bin packing problem under uniform and Weibull distributions. In this work, we reassess this claim through a detailed analysis of the heuristics produced by LLMs, examining both their beha
An In-depth Study of LLM Contributions to the Bin Packing Problem
-
cs.AI, q-bio.NC updates on arXiv.org
-
VeriMoA: A Mixture-of-Agents Framework for Spec-to-HDL Generation
arXiv:2510.27617v1 Announce Type: new Abstract: Automation of Register Transfer Level (RTL) design can help developers meet increasing computational demands. Large Language Models (LLMs) show promise for Hardware Description Language (HDL) generation, but face challenges due to limited parametric knowledge and domain-specific constraints. While prompt engineering and fine-tuning have limitations in knowledge coverage and training costs, multi-agent architectures offer a training-free paradigm t
VeriMoA: A Mixture-of-Agents Framework for Spec-to-HDL Generation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Impact of clinical decision support systems (cdss) on clinical outcomes and healthcare delivery in low- and middle-income countries: protocol for a systematic review and meta-analysis
arXiv:2510.26812v1 Announce Type: cross Abstract: Clinical decision support systems (CDSS) are used to improve clinical and service outcomes, yet evidence from low- and middle-income countries (LMICs) is dispersed. This protocol outlines methods to quantify the impact of CDSS on patient and healthcare delivery outcomes in LMICs. We will include comparative quantitative designs (randomized trials, controlled before-after, interrupted time series, comparative cohorts) evaluating CDSS in World Ban
Impact of clinical decision support systems (cdss) on clinical outcomes and healthcare delivery in low- and middle-income countries: protocol for a systematic review and meta-analysis
-
cs.AI, q-bio.NC updates on arXiv.org
-
Frame Semantic Patterns for Identifying Underreporting of Notifiable Events in Healthcare: The Case of Gender-Based Violence
arXiv:2510.26969v1 Announce Type: cross Abstract: We introduce a methodology for the identification of notifiable events in the domain of healthcare. The methodology harnesses semantic frames to define fine-grained patterns and search them in unstructured data, namely, open-text fields in e-medical records. We apply the methodology to the problem of underreporting of gender-based violence (GBV) in e-medical records produced during patients' visits to primary care units. A total of eight pattern
Frame Semantic Patterns for Identifying Underreporting of Notifiable Events in Healthcare: The Case of Gender-Based Violence
-
cs.AI, q-bio.NC updates on arXiv.org
-
MARIA: A Framework for Marginal Risk Assessment without Ground Truth in AI Systems
arXiv:2510.27163v1 Announce Type: cross Abstract: Before deploying an AI system to replace an existing process, it must be compared with the incumbent to ensure improvement without added risk. Traditional evaluation relies on ground truth for both systems, but this is often unavailable due to delayed or unknowable outcomes, high costs, or incomplete data, especially for long-standing systems deemed safe by convention. The more practical solution is not to compute absolute risk but the differenc
MARIA: A Framework for Marginal Risk Assessment without Ground Truth in AI Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
arXiv:2510.27196v1 Announce Type: cross Abstract: The proliferation of memes on social media necessitates the capabilities of multimodal Large Language Models (mLLMs) to effectively understand multimodal harmfulness. Existing evaluation approaches predominantly focus on mLLMs' detection accuracy for binary classification tasks, which often fail to reflect the in-depth interpretive nuance of harmfulness across diverse contexts. In this paper, we propose MemeArena, an agent-based arena-style eval
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
arXiv:2510.27287v1 Announce Type: cross Abstract: Enterprise systems are crucial for enhancing productivity and decision-making among employees and customers. Integrating LLM based systems into enterprise systems enables intelligent automation, personalized experiences, and efficient information retrieval, driving operational efficiency and strategic growth. However, developing and evaluating such systems is challenging due to the inherent complexity of enterprise environments, where data is fr
Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
-
cs.AI, q-bio.NC updates on arXiv.org
-
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
arXiv:2510.27629v1 Announce Type: cross Abstract: Open-weight bio-foundation models present a dual-use dilemma. While holding great promise for accelerating scientific research and drug development, they could also enable bad actors to develop more deadly bioweapons. To mitigate the risk posed by these models, current approaches focus on filtering biohazardous data during pre-training. However, the effectiveness of such an approach remains unclear, particularly against determined actors who mig
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Systematic Literature Review of Spatio-Temporal Graph Neural Network Models for Time Series Forecasting and Classification
arXiv:2410.22377v3 Announce Type: replace-cross Abstract: In recent years, spatio-temporal graph neural networks (GNNs) have attracted considerable interest in the field of time series analysis, due to their ability to capture, at once, dependencies among variables and across time points. The objective of this systematic literature review is hence to provide a comprehensive overview of the various modeling approaches and application domains of GNNs for time series classification and forecasting