Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Denario project: Deep knowledge AI agents for scientific discovery
arXiv:2510.26887v1 Announce Type: new Abstract: We present Denario, an AI multi-agent system designed to serve as a scientific research assistant. Denario can perform many different tasks, such as generating ideas, checking the literature, developing research plans, writing and executing code, making plots, and drafting and reviewing a scientific paper. The system has a modular architecture, allowing it to handle specific tasks, such as generating an idea, or carrying out end-to-end scientific
-
cs.AI, q-bio.NC updates on arXiv.org
-
Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement
arXiv:2510.27051v1 Announce Type: new Abstract: Enterprise AI agents must continuously adapt to maintain accuracy, reduce latency, and remain aligned with user needs. We present a practical implementation of a data flywheel in NVInfo AI, NVIDIA's Mixture-of-Experts (MoE) Knowledge Assistant serving over 30,000 employees. By operationalizing a MAPE-driven data flywheel, we built a closed-loop system that systematically addresses failures in retrieval-augmented generation (RAG) pipelines and enab
Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement
-
cs.AI, q-bio.NC updates on arXiv.org
-
Glia: A Human-Inspired AI for Automated Systems Design and Optimization
arXiv:2510.27176v1 Announce Type: new Abstract: Can an AI autonomously design mechanisms for computer systems on par with the creativity and reasoning of human experts? We present Glia, an AI architecture for networked systems design that uses large language models (LLMs) in a human-inspired, multi-agent workflow. Each agent specializes in reasoning, experimentation, and analysis, collaborating through an evaluation framework that grounds abstract reasoning in empirical feedback. Unlike prior M
Glia: A Human-Inspired AI for Automated Systems Design and Optimization
-
cs.AI, q-bio.NC updates on arXiv.org
-
An In-depth Study of LLM Contributions to the Bin Packing Problem
arXiv:2510.27353v1 Announce Type: new Abstract: Recent studies have suggested that Large Language Models (LLMs) could provide interesting ideas contributing to mathematical discovery. This claim was motivated by reports that LLM-based genetic algorithms produced heuristics offering new insights into the online bin packing problem under uniform and Weibull distributions. In this work, we reassess this claim through a detailed analysis of the heuristics produced by LLMs, examining both their beha
An In-depth Study of LLM Contributions to the Bin Packing Problem
-
cs.AI, q-bio.NC updates on arXiv.org
-
VeriMoA: A Mixture-of-Agents Framework for Spec-to-HDL Generation
arXiv:2510.27617v1 Announce Type: new Abstract: Automation of Register Transfer Level (RTL) design can help developers meet increasing computational demands. Large Language Models (LLMs) show promise for Hardware Description Language (HDL) generation, but face challenges due to limited parametric knowledge and domain-specific constraints. While prompt engineering and fine-tuning have limitations in knowledge coverage and training costs, multi-agent architectures offer a training-free paradigm t
VeriMoA: A Mixture-of-Agents Framework for Spec-to-HDL Generation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Impact of clinical decision support systems (cdss) on clinical outcomes and healthcare delivery in low- and middle-income countries: protocol for a systematic review and meta-analysis
arXiv:2510.26812v1 Announce Type: cross Abstract: Clinical decision support systems (CDSS) are used to improve clinical and service outcomes, yet evidence from low- and middle-income countries (LMICs) is dispersed. This protocol outlines methods to quantify the impact of CDSS on patient and healthcare delivery outcomes in LMICs. We will include comparative quantitative designs (randomized trials, controlled before-after, interrupted time series, comparative cohorts) evaluating CDSS in World Ban
Impact of clinical decision support systems (cdss) on clinical outcomes and healthcare delivery in low- and middle-income countries: protocol for a systematic review and meta-analysis
-
cs.AI, q-bio.NC updates on arXiv.org
-
Frame Semantic Patterns for Identifying Underreporting of Notifiable Events in Healthcare: The Case of Gender-Based Violence
arXiv:2510.26969v1 Announce Type: cross Abstract: We introduce a methodology for the identification of notifiable events in the domain of healthcare. The methodology harnesses semantic frames to define fine-grained patterns and search them in unstructured data, namely, open-text fields in e-medical records. We apply the methodology to the problem of underreporting of gender-based violence (GBV) in e-medical records produced during patients' visits to primary care units. A total of eight pattern
Frame Semantic Patterns for Identifying Underreporting of Notifiable Events in Healthcare: The Case of Gender-Based Violence
-
cs.AI, q-bio.NC updates on arXiv.org
-
MARIA: A Framework for Marginal Risk Assessment without Ground Truth in AI Systems
arXiv:2510.27163v1 Announce Type: cross Abstract: Before deploying an AI system to replace an existing process, it must be compared with the incumbent to ensure improvement without added risk. Traditional evaluation relies on ground truth for both systems, but this is often unavailable due to delayed or unknowable outcomes, high costs, or incomplete data, especially for long-standing systems deemed safe by convention. The more practical solution is not to compute absolute risk but the differenc
MARIA: A Framework for Marginal Risk Assessment without Ground Truth in AI Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
arXiv:2510.27196v1 Announce Type: cross Abstract: The proliferation of memes on social media necessitates the capabilities of multimodal Large Language Models (mLLMs) to effectively understand multimodal harmfulness. Existing evaluation approaches predominantly focus on mLLMs' detection accuracy for binary classification tasks, which often fail to reflect the in-depth interpretive nuance of harmfulness across diverse contexts. In this paper, we propose MemeArena, an agent-based arena-style eval
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
arXiv:2510.27287v1 Announce Type: cross Abstract: Enterprise systems are crucial for enhancing productivity and decision-making among employees and customers. Integrating LLM based systems into enterprise systems enables intelligent automation, personalized experiences, and efficient information retrieval, driving operational efficiency and strategic growth. However, developing and evaluating such systems is challenging due to the inherent complexity of enterprise environments, where data is fr
Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
-
cs.AI, q-bio.NC updates on arXiv.org
-
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
arXiv:2510.27629v1 Announce Type: cross Abstract: Open-weight bio-foundation models present a dual-use dilemma. While holding great promise for accelerating scientific research and drug development, they could also enable bad actors to develop more deadly bioweapons. To mitigate the risk posed by these models, current approaches focus on filtering biohazardous data during pre-training. However, the effectiveness of such an approach remains unclear, particularly against determined actors who mig
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Systematic Literature Review of Spatio-Temporal Graph Neural Network Models for Time Series Forecasting and Classification
arXiv:2410.22377v3 Announce Type: replace-cross Abstract: In recent years, spatio-temporal graph neural networks (GNNs) have attracted considerable interest in the field of time series analysis, due to their ability to capture, at once, dependencies among variables and across time points. The objective of this systematic literature review is hence to provide a comprehensive overview of the various modeling approaches and application domains of GNNs for time series classification and forecasting
A Systematic Literature Review of Spatio-Temporal Graph Neural Network Models for Time Series Forecasting and Classification
-
cs.AI, q-bio.NC updates on arXiv.org
-
Artificial Empathy: AI based Mental Health
arXiv:2506.00081v2 Announce Type: replace-cross Abstract: Many people suffer from mental health problems but not everyone seeks professional help or has access to mental health care. AI chatbots have increasingly become a go-to for individuals who either have mental disorders or simply want someone to talk to. This paper presents a study on participants who have previously used chatbots and a scenario-based testing of large language model (LLM) chatbots. Our findings indicate that AI chatbots w
Artificial Empathy: AI based Mental Health
-
cs.AI, q-bio.NC updates on arXiv.org
-
Deep Learning-based Prediction of Clinical Trial Enrollment with Uncertainty Estimates
arXiv:2507.23607v2 Announce Type: replace-cross Abstract: Clinical trials are a systematic endeavor to assess the safety and efficacy of new drugs or treatments. Conducting such trials typically demands significant financial investment and meticulous planning, highlighting the need for accurate predictions of trial outcomes. Accurately predicting patient enrollment, a key factor in trial success, is one of the primary challenges during the planning phase. In this work, we propose a novel deep l
Deep Learning-based Prediction of Clinical Trial Enrollment with Uncertainty Estimates
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Process Mining-Based System For The Analysis and Prediction of Software Development Workflows
arXiv:2510.25935v2 Announce Type: replace-cross Abstract: CodeSight is an end-to-end system designed to anticipate deadline compliance in software development workflows. It captures development and deployment data directly from GitHub, transforming it into process mining logs for detailed analysis. From these logs, the system generates metrics and dashboards that provide actionable insights into PR activity patterns and workflow efficiency. Building on this structured representation, CodeSight
A Process Mining-Based System For The Analysis and Prediction of Software Development Workflows
-
cs.AI, q-bio.NC updates on arXiv.org
-
On the limitation of evaluating machine unlearning using only a single training seed
arXiv:2510.26714v2 Announce Type: replace-cross Abstract: Machine unlearning (MU) aims to remove the influence of certain data points from a trained model without costly retraining. Most practical MU algorithms are only approximate and their performance can only be assessed empirically. Care must therefore be taken to make empirical comparisons as representative as possible. A common practice is to run the MU algorithm multiple times independently starting from the same trained model. In this w
On the limitation of evaluating machine unlearning using only a single training seed
-
MRD
-
International expert consensus on the clinical integration of circulating tumor cells in solid tumors
Eur J Cancer. 2025 Dec 9;231:116050. doi: 10.1016/j.ejca.2025.116050. Epub 2025 Oct 20.ABSTRACTBACKGROUND: Circulating tumor cells (CTCs) are a versatile biomarker in solid tumors. Extensive research supports their clinical relevance and led to regulatory approval in breast, prostate, and colorectal cancers. However, clinical adoption remains limited mainly due to the lack of consensus and standardized technologies. Additionally, CTC research lacks unified direction. To address these gaps, an in
International expert consensus on the clinical integration of circulating tumor cells in solid tumors
Eur J Cancer. 2025 Dec 9;231:116050. doi: 10.1016/j.ejca.2025.116050. Epub 2025 Oct 20.
ABSTRACT
BACKGROUND: Circulating tumor cells (CTCs) are a versatile biomarker in solid tumors. Extensive research supports their clinical relevance and led to regulatory approval in breast, prostate, and colorectal cancers. However, clinical adoption remains limited mainly due to the lack of consensus and standardized technologies. Additionally, CTC research lacks unified direction. To address these gaps, an international expert panel was established to assess the current and future clinical utility of CTCs.
METHODS: A panel of 11 CTC experts identified key areas of controversy, informing a structured survey distributed to 55 international multidisciplinary experts. Consensus was predefined as ≥ 70 % agreement. Areas without consensus were discussed in a virtual meeting, leading to final statements on the clinical integration of CTCs.
RESULTS: Thirty-seven experts completed the survey. Consensus was reached on the clinical utility of CTCs for prognosis and treatment monitoring in metastatic breast (BC) and prostate (PC) cancers, including AR-V7 testing in metastatic castration-resistant PC for therapy selection. In other tumors, CTCs remain investigational. Experts agreed that while clinical utility is not yet established in early-stage disease, CTCs show promise in early BC, especially combined with cell-free DNA (cfDNA) for minimal residual disease detection. CellSearch® is currently the only platform with high-level evidence for clinical use, though emerging technologies are promising. Key challenges include improving detection sensitivity/specificity, standardizing workflows, generating robust data, and clinician education. Experts emphasized shifting from enumeration to phenotypic and molecular characterization, particularly for treatment guidance, and highlighted the complementary role of CTCs and cfDNA, advocating for integrated liquid biopsy approaches.
CONCLUSIONS: This consensus offers practical guidance for clinical integration of CTCs and outlines strategic research priorities to unlock their full potential in precision oncology.
PMID:41172567 | DOI:10.1016/j.ejca.2025.116050
-
(Multiomics OR Omics) AND (Lung OR gastric OR Hepatocellular)
-
Animal models in tuberculosis metabolomics: a systematic review of current evidence and the road to translational relevance
Front Mol Biosci. 2025 Oct 15;12:1688882. doi: 10.3389/fmolb.2025.1688882. eCollection 2025.ABSTRACTBACKGROUND: Animal models are important for tuberculosis (TB) research, offering controlled settings to study disease mechanisms. However, their ability to replicate TB-induced metabolic responses in humans is uncertain. This systematic review evaluated the current use of animal models in metabolomics studies aimed at characterising active pulmonary TB.METHODS: PubMed, Scopus, and Web of Science w
Animal models in tuberculosis metabolomics: a systematic review of current evidence and the road to translational relevance
Front Mol Biosci. 2025 Oct 15;12:1688882. doi: 10.3389/fmolb.2025.1688882. eCollection 2025.
ABSTRACT
BACKGROUND: Animal models are important for tuberculosis (TB) research, offering controlled settings to study disease mechanisms. However, their ability to replicate TB-induced metabolic responses in humans is uncertain. This systematic review evaluated the current use of animal models in metabolomics studies aimed at characterising active pulmonary TB.
METHODS: PubMed, Scopus, and Web of Science were systematically searched for metabolomics studies of pulmonary TB in humans and animal models, following PRISMA guidelines. Eligible studies were screened, and quality was assessed using QUDOMICS and STAIR tools. Data were synthesised by species, sample matrix, experimental design, and reported differential metabolites. Differential metabolite names were compared between species and subjected to pathway analysis in MetaboAnalyst 6.0.
RESULTS: Of the 80 eligible studies, nine involved animal models, predominantly mice. These models captured only 4.7% of human TB-associated differential metabolites, with the highest overlap (3.8%) in mouse lung tissue. Despite low concordance at metabolite level, conserved disruptions were observed in amino acid, glutathione, and one-carbon metabolism pathways. Interspecies variation was evident, influenced by host species, sample matrix, infection protocol, and analytical method.
CONCLUSION: Animal models partially replicated key metabolic features of human TB, particularly at the pathway level. However, variability across studies hampers current translational interpretation. Broader model use, standardised protocols, and integrated multi-platform omics approaches are needed to improve the relevance and comparability of animal models in TB metabolomics research.
PMID:41169614 | PMC:PMC12568366 | DOI:10.3389/fmolb.2025.1688882
-
Journal of Medical Internet Research
-
Wearable Artificial Intelligence for Epilepsy: Scoping Review
Background: Epilepsy affects approximately 50 million people globally and imposes a substantial clinical and societal burden, requiring continuous and personalized monitoring for effective management. Wearable artificial intelligence (AI) technologies offer a promising solution by leveraging physiological signals and machine learning for seizure detection and prediction. While various approaches have been proposed, a comprehensive overview summarizing these advances and challenges is still neede
Wearable Artificial Intelligence for Epilepsy: Scoping Review
-
Journal of Medical Internet Research
-
Opportunities and Challenges for Designing in Connected Health: Insights From an Expert Workshop
Health care increasingly depends on information and communication technology. This offers both opportunities and challenges when designing connected health systems. While individual studies examined particular cases, there is a limited synthesis of insights across projects. The objective of this paper is to explore these opportunities and challenges by examining 6 diverse connected health projects and synthesizing lessons from an expert workshop. To achieve this, we conducted a full-day workshop