❌

Reading view

How can we assess human-agent interactions? Case studies in software agent design

arXiv:2510.09801v2 Announce Type: replace Abstract: LLM-powered agents are both a promising new technology and a source of complexity, where choices about models, tools, and prompting can affect their usefulness. While numerous benchmarks measure agent accuracy across domains, they mostly assume full automation, failing to represent the collaborative nature of real-world use cases. In this paper, we make two major steps towards the rigorous assessment of human-agent interactions. First, we propose PULSE, a framework for more efficient human-centric evaluation of agent designs, which comprises collecting user feedback, training an ML model to predict user satisfaction, and computing results by combining human satisfaction ratings with model-generated pseudo-labels. Second, we deploy the framework on a large-scale web platform built around the open-source software agent OpenHands, collecting in-the-wild usage data across over 15k users. We conduct case studies around how three agent design decisions -- choice of LLM backbone, planning strategy, and memory mechanisms -- impact developer satisfaction rates, yielding practical insights for software agent design. We also show how our framework can lead to more robust conclusions about agent design, reducing confidence intervals by 40% compared to a standard A/B test. Finally, we find substantial discrepancies between in-the-wild results and benchmark performance (e.g., the anti-correlation between results comparing claude-sonnet-4 and gpt-5), underscoring the limitations of benchmark-driven evaluation. Our findings provide guidance for evaluations of LLM agents with humans and identify opportunities for better agent designs.
  •  

Digital Twin based Automatic Reconfiguration of Robotic Systems in Smart Environments

arXiv:2511.00094v1 Announce Type: cross Abstract: Robotic systems have become integral to smart environments, enabling applications ranging from urban surveillance and automated agriculture to industrial automation. However, their effective operation in dynamic settings - such as smart cities and precision farming - is challenged by continuously evolving topographies and environmental conditions. Traditional control systems often struggle to adapt quickly, leading to inefficiencies or operational failures. To address this limitation, we propose a novel framework for autonomous and dynamic reconfiguration of robotic controllers using Digital Twin technology. Our approach leverages a virtual replica of the robot's operational environment to simulate and optimize movement trajectories in response to real-world changes. By recalculating paths and control parameters in the Digital Twin and deploying the updated code to the physical robot, our method ensures rapid and reliable adaptation without manual intervention. This work advances the integration of Digital Twins in robotics, offering a scalable solution for enhancing autonomy in smart, dynamic environments.
  •  

Will Humanity Be Rendered Obsolete by AI?

arXiv:2510.22814v2 Announce Type: replace Abstract: This article analyzes the existential risks artificial intelligence (AI) poses to humanity, tracing the trajectory from current AI to ultraintelligence. Drawing on Irving J. Good and Nick Bostrom's theoretical work, plus recent publications (AI 2027; If Anyone Builds It, Everyone Dies), it explores AGI and superintelligence. Considering machines' exponentially growing cognitive power and hypothetical IQs, it addresses the ethical and existential implications of an intelligence vastly exceeding humanity's, fundamentally alien. Human extinction may result not from malice, but from uncontrollable, indifferent cognitive superiority.
  •  

The Denario project: Deep knowledge AI agents for scientific discovery

arXiv:2510.26887v1 Announce Type: new Abstract: We present Denario, an AI multi-agent system designed to serve as a scientific research assistant. Denario can perform many different tasks, such as generating ideas, checking the literature, developing research plans, writing and executing code, making plots, and drafting and reviewing a scientific paper. The system has a modular architecture, allowing it to handle specific tasks, such as generating an idea, or carrying out end-to-end scientific analysis using Cmbagent as a deep-research backend. In this work, we describe in detail Denario and its modules, and illustrate its capabilities by presenting multiple AI-generated papers generated by it in many different scientific disciplines such as astrophysics, biology, biophysics, biomedical informatics, chemistry, material science, mathematical physics, medicine, neuroscience and planetary science. Denario also excels at combining ideas from different disciplines, and we illustrate this by showing a paper that applies methods from quantum physics and machine learning to astrophysical data. We report the evaluations performed on these papers by domain experts, who provided both numerical scores and review-like feedback. We then highlight the strengths, weaknesses, and limitations of the current system. Finally, we discuss the ethical implications of AI-driven research and reflect on how such technology relates to the philosophy of science. We publicly release the code at https://github.com/AstroPilot-AI/Denario. A Denario demo can also be run directly on the web at https://huggingface.co/spaces/astropilot-ai/Denario, and the full app will be deployed on the cloud.
  •  

Frame Semantic Patterns for Identifying Underreporting of Notifiable Events in Healthcare: The Case of Gender-Based Violence

arXiv:2510.26969v1 Announce Type: cross Abstract: We introduce a methodology for the identification of notifiable events in the domain of healthcare. The methodology harnesses semantic frames to define fine-grained patterns and search them in unstructured data, namely, open-text fields in e-medical records. We apply the methodology to the problem of underreporting of gender-based violence (GBV) in e-medical records produced during patients' visits to primary care units. A total of eight patterns are defined and searched on a corpus of 21 million sentences in Brazilian Portuguese extracted from e-SUS APS. The results are manually evaluated by linguists and the precision of each pattern measured. Our findings reveal that the methodology effectively identifies reports of violence with a precision of 0.726, confirming its robustness. Designed as a transparent, efficient, low-carbon, and language-agnostic pipeline, the approach can be easily adapted to other health surveillance contexts, contributing to the broader, ethical, and explainable use of NLP in public health systems.
  •  

A Systematic Literature Review of Spatio-Temporal Graph Neural Network Models for Time Series Forecasting and Classification

arXiv:2410.22377v3 Announce Type: replace-cross Abstract: In recent years, spatio-temporal graph neural networks (GNNs) have attracted considerable interest in the field of time series analysis, due to their ability to capture, at once, dependencies among variables and across time points. The objective of this systematic literature review is hence to provide a comprehensive overview of the various modeling approaches and application domains of GNNs for time series classification and forecasting. A database search was conducted, and 366 papers were selected for a detailed examination of the current state-of-the-art in the field. This examination is intended to offer to the reader a comprehensive review of proposed models, links to related source code, available datasets, benchmark models, and fitting results. All this information is hoped to assist researchers in their studies. To the best of our knowledge, this is the first and broadest systematic literature review presenting a detailed comparison of results from current spatio-temporal GNN models applied to different domains. In its final part, this review discusses current limitations and challenges in the application of spatio-temporal GNNs, such as comparability, reproducibility, explainability, poor information capacity, and scalability. This paper is complemented by a GitHub repository at https://github.com/FlaGer99/SLR-Spatio-Temporal-GNN.git providing additional interactive tools to further explore the presented findings.
  •  

Deep Learning-based Prediction of Clinical Trial Enrollment with Uncertainty Estimates

arXiv:2507.23607v2 Announce Type: replace-cross Abstract: Clinical trials are a systematic endeavor to assess the safety and efficacy of new drugs or treatments. Conducting such trials typically demands significant financial investment and meticulous planning, highlighting the need for accurate predictions of trial outcomes. Accurately predicting patient enrollment, a key factor in trial success, is one of the primary challenges during the planning phase. In this work, we propose a novel deep learning-based method to address this critical challenge. Our method, implemented as a neural network model, leverages pre-trained language models (PLMs) to capture the complexities and nuances of clinical documents, transforming them into expressive representations. These representations are then combined with encoded tabular features via an attention mechanism. To account for uncertainties in enrollment prediction, we enhance the model with a probabilistic layer based on the Gamma distribution, which enables range estimation. We apply the proposed model to predict clinical trial duration, assuming site-level enrollment follows a Poisson-Gamma process. We carry out extensive experiments on real-world clinical trial data, and show that the proposed method can effectively predict the number of patients enrolled at a number of sites for a given clinical trial, outperforming established baseline models.
  •  

A Process Mining-Based System For The Analysis and Prediction of Software Development Workflows

arXiv:2510.25935v2 Announce Type: replace-cross Abstract: CodeSight is an end-to-end system designed to anticipate deadline compliance in software development workflows. It captures development and deployment data directly from GitHub, transforming it into process mining logs for detailed analysis. From these logs, the system generates metrics and dashboards that provide actionable insights into PR activity patterns and workflow efficiency. Building on this structured representation, CodeSight employs an LSTM model that predicts remaining PR resolution times based on sequential activity traces and static features, enabling early identification of potential deadline breaches. In tests, the system demonstrates high precision and F1 scores in predicting deadline compliance, illustrating the value of integrating process mining with machine learning for proactive software project management.
  •  

On the limitation of evaluating machine unlearning using only a single training seed

arXiv:2510.26714v2 Announce Type: replace-cross Abstract: Machine unlearning (MU) aims to remove the influence of certain data points from a trained model without costly retraining. Most practical MU algorithms are only approximate and their performance can only be assessed empirically. Care must therefore be taken to make empirical comparisons as representative as possible. A common practice is to run the MU algorithm multiple times independently starting from the same trained model. In this work, we demonstrate that this practice can give highly non-representative results because -- even for the same architecture and same dataset -- some MU methods can be highly sensitive to the choice of random number seed used for model training. We therefore recommend that empirical comparisons of MU algorithms should also reflect the variability across different model training seeds.
  •  

International expert consensus on the clinical integration of circulating tumor cells in solid tumors

Eur J Cancer. 2025 Dec 9;231:116050. doi: 10.1016/j.ejca.2025.116050. Epub 2025 Oct 20.

ABSTRACT

BACKGROUND: Circulating tumor cells (CTCs) are a versatile biomarker in solid tumors. Extensive research supports their clinical relevance and led to regulatory approval in breast, prostate, and colorectal cancers. However, clinical adoption remains limited mainly due to the lack of consensus and standardized technologies. Additionally, CTC research lacks unified direction. To address these gaps, an international expert panel was established to assess the current and future clinical utility of CTCs.

METHODS: A panel of 11 CTC experts identified key areas of controversy, informing a structured survey distributed to 55 international multidisciplinary experts. Consensus was predefined as ≥ 70 % agreement. Areas without consensus were discussed in a virtual meeting, leading to final statements on the clinical integration of CTCs.

RESULTS: Thirty-seven experts completed the survey. Consensus was reached on the clinical utility of CTCs for prognosis and treatment monitoring in metastatic breast (BC) and prostate (PC) cancers, including AR-V7 testing in metastatic castration-resistant PC for therapy selection. In other tumors, CTCs remain investigational. Experts agreed that while clinical utility is not yet established in early-stage disease, CTCs show promise in early BC, especially combined with cell-free DNA (cfDNA) for minimal residual disease detection. CellSearch® is currently the only platform with high-level evidence for clinical use, though emerging technologies are promising. Key challenges include improving detection sensitivity/specificity, standardizing workflows, generating robust data, and clinician education. Experts emphasized shifting from enumeration to phenotypic and molecular characterization, particularly for treatment guidance, and highlighted the complementary role of CTCs and cfDNA, advocating for integrated liquid biopsy approaches.

CONCLUSIONS: This consensus offers practical guidance for clinical integration of CTCs and outlines strategic research priorities to unlock their full potential in precision oncology.

PMID:41172567 | DOI:10.1016/j.ejca.2025.116050

  •  

Animal models in tuberculosis metabolomics: a systematic review of current evidence and the road to translational relevance

Front Mol Biosci. 2025 Oct 15;12:1688882. doi: 10.3389/fmolb.2025.1688882. eCollection 2025.

ABSTRACT

BACKGROUND: Animal models are important for tuberculosis (TB) research, offering controlled settings to study disease mechanisms. However, their ability to replicate TB-induced metabolic responses in humans is uncertain. This systematic review evaluated the current use of animal models in metabolomics studies aimed at characterising active pulmonary TB.

METHODS: PubMed, Scopus, and Web of Science were systematically searched for metabolomics studies of pulmonary TB in humans and animal models, following PRISMA guidelines. Eligible studies were screened, and quality was assessed using QUDOMICS and STAIR tools. Data were synthesised by species, sample matrix, experimental design, and reported differential metabolites. Differential metabolite names were compared between species and subjected to pathway analysis in MetaboAnalyst 6.0.

RESULTS: Of the 80 eligible studies, nine involved animal models, predominantly mice. These models captured only 4.7% of human TB-associated differential metabolites, with the highest overlap (3.8%) in mouse lung tissue. Despite low concordance at metabolite level, conserved disruptions were observed in amino acid, glutathione, and one-carbon metabolism pathways. Interspecies variation was evident, influenced by host species, sample matrix, infection protocol, and analytical method.

CONCLUSION: Animal models partially replicated key metabolic features of human TB, particularly at the pathway level. However, variability across studies hampers current translational interpretation. Broader model use, standardised protocols, and integrated multi-platform omics approaches are needed to improve the relevance and comparability of animal models in TB metabolomics research.

PMID:41169614 | PMC:PMC12568366 | DOI:10.3389/fmolb.2025.1688882

  •  

Digital Health Technology Compliance With Clinical Safety Standards In the National Health Service in England: National Cross-Sectional Study

Background: To be authorized for use in the National Health Service (NHS) in England, digital health technologies (DHTs) must meet 2 mandatory clinical risk management standards, Data Coordination Board (DCB) 0129 and 0160, demonstrating that risks from design and use have been assessed and mitigated. NHS organizations must not procure a DHT without DCB0129 assurance and must not deploy one without DCB0160 assurance. Despite legal requirement, no public data exist on how many DHTs are in use in the NHS or how many are assured. Objective: This study aimed to determine the number of DHTs in use in the NHS in England and assess their assurance status against mandated clinical safety standards. Methods: In early 2025, 239 NHS organizations in England received a freedom of information notice requesting information on the number of DHTs they were using and their assurance against DCB0129 and DCB0160 standards. Results: Of the 239 NHS organizations, 204 (85.4%) responded, of which 178 (87.3%) provided full or partial data, covering 14,747 DHT deployments. The mean number of deployed DHTs per organization was 82.8 (SD 146.1; 95% CI 61.4-104.3) with substantial variation between NHS provider trusts (mean 107.1, SD 161.1; 95% CI 79.8-134.3), ambulance trusts (mean 13.0, SD 8.2; 95% CI 7.6-18.4), and integrated care boards (mean 8.1, SD 16.0; 95% CI 2.8-13.5). Overall organizational compliance rates were low, with a median of 25.6% (IQR 7.8%-55.7%) deployed DHTs being fully assured; for NHS provider trusts compliance was lower at 24.5% (IQR 8.1%-50%). A total of 13 (6.4%) of the 204 organizations reported that all their DHTs were fully assured, while 16 (7.8%) reported that none were assured. Across all DHTs with reported assurance data, 17.3% (95% CI 16.6%-18.1%) were fully assured against both standards, 13.3% were partially assured against one standard, and 70.1% (95% CI 69.1%-71.1%) had no documented assurance. Conclusions: This is the first study to quantify both the scale of DHT deployment in NHS organizations in England and the extent of compliance with mandatory safety standards. More than 10,000 DHTs currently in use lack documented assurance against clinical safety standards. In a typical NHS trust, 3 out of 4 digital tools influencing patient care do not demonstrate compliance with minimum legal or clinical safety requirements. These findings raise significant concerns about the risks posed to patients by these technologies; the capacity of organizations to assess and mitigate them; and the legal ramifications of when, not if, harm occurs. Crucially, failure to assure digital technologies poses a significant risk to one of the core ambitions of the NHS 10-Year Health Plan for England; safely transitioning from analogue to digital care models. These findings are unlikely to be unique to the NHS and should prompt health care systems worldwide to assess the risks posed by their DHT deployments.
  •  

Opportunities and Challenges for Designing in Connected Health: Insights From an Expert Workshop

Health care increasingly depends on information and communication technology. This offers both opportunities and challenges when designing connected health systems. While individual studies examined particular cases, there is a limited synthesis of insights across projects. The objective of this paper is to explore these opportunities and challenges by examining 6 diverse connected health projects and synthesizing lessons from an expert workshop. To achieve this, we conducted a full-day workshop that brought together 6 connected health projects. The workshop used an iterative and participatory process which included paper submissions and presentations and facilitated discussions, and a gallery walk to enable cross-case comparison and collaborative reflection. Thematic analysis of workshop outputs was then used to synthesize key opportunities and challenges in designing connected health systems. The 6 projects represented a variety of design methods and approaches to connected health, and their discussion surfaced both opportunities and challenges in this domain. Key opportunities include improving data integration and usability, enhancing collaboration across stakeholders, using a user-centered and iterative design process, addressing complexity in sociotechnical systems, sustainability, and adopting digital infrastructures for seamless communication. Participants also identified important challenges, namely exchange of information, interoperability, and communication; ethical considerations, rules, and regulations; understanding design, evaluation, and standards; actionable data, reliability, quality, and trust in data; and stakeholder involvement. The contribution of this paper lies in the synthesis of insights across multiple projects and perspectives to provide practical guidance for researchers, designers, and policymakers. By highlighting opportunities and challenges in designing connected health systems, the findings emphasize the importance of patient-centered, sustainable, and collaborative design approaches while also pointing to the need to address persistent barriers. Advancing connected health will require adopting iterative and inclusive design processes that prioritize patient-centeredness, sustainability, and collaboration across health care systems.
  •  

An Agentic Framework for Rapid Deployment of Edge AI Solutions in Industry 5.0

arXiv:2510.25813v1 Announce Type: new Abstract: We present a novel framework for Industry 5.0 that simplifies the deployment of AI models on edge devices in various industrial settings. The design reduces latency and avoids external data transfer by enabling local inference and real-time processing. Our implementation is agent-based, which means that individual agents, whether human, algorithmic, or collaborative, are responsible for well-defined tasks, enabling flexibility and simplifying integration. Moreover, our framework supports modular integration and maintains low resource requirements. Preliminary evaluations concerning the food industry in real scenarios indicate improved deployment time and system adaptability performance. The source code is publicly available at https://github.com/AI-REDGIO-5-0/ci-component.
  •  

SciTrust 2.0: A Comprehensive Framework for Evaluating Trustworthiness of Large Language Models in Scientific Applications

arXiv:2510.25908v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated transformative potential in scientific research, yet their deployment in high-stakes contexts raises significant trustworthiness concerns. Here, we introduce SciTrust 2.0, a comprehensive framework for evaluating LLM trustworthiness in scientific applications across four dimensions: truthfulness, adversarial robustness, scientific safety, and scientific ethics. Our framework incorporates novel, open-ended truthfulness benchmarks developed through a verified reflection-tuning pipeline and expert validation, alongside a novel ethics benchmark for scientific research contexts covering eight subcategories including dual-use research and bias. We evaluated seven prominent LLMs, including four science-specialized models and three general-purpose industry models, using multiple evaluation metrics including accuracy, semantic similarity measures, and LLM-based scoring. General-purpose industry models overall outperformed science-specialized models across each trustworthiness dimension, with GPT-o4-mini demonstrating superior performance in truthfulness assessments and adversarial robustness. Science-specialized models showed significant deficiencies in logical and ethical reasoning capabilities, along with concerning vulnerabilities in safety evaluations, particularly in high-risk domains such as biosecurity and chemical weapons. By open-sourcing our framework, we provide a foundation for developing more trustworthy AI systems and advancing research on model safety and ethics in scientific contexts.
  •  

Human-AI Complementarity: A Goal for Amplified Oversight

arXiv:2510.26518v1 Announce Type: new Abstract: Human feedback is critical for aligning AI systems to human values. As AI capabilities improve and AI is used to tackle more challenging tasks, verifying quality and safety becomes increasingly challenging. This paper explores how we can leverage AI to improve the quality of human oversight. We focus on an important safety problem that is already challenging for humans: fact-verification of AI outputs. We find that combining AI ratings and human ratings based on AI rater confidence is better than relying on either alone. Giving humans an AI fact-verification assistant further improves their accuracy, but the type of assistance matters. Displaying AI explanation, confidence, and labels leads to over-reliance, but just showing search results and evidence fosters more appropriate trust. These results have implications for Amplified Oversight -- the challenge of combining humans and AI to supervise AI systems even as they surpass human expert performance.
  •  

Agentic AI Home Energy Management System: A Large Language Model Framework for Residential Load Scheduling

arXiv:2510.26603v1 Announce Type: new Abstract: The electricity sector transition requires substantial increases in residential demand response capacity, yet Home Energy Management Systems (HEMS) adoption remains limited by user interaction barriers requiring translation of everyday preferences into technical parameters. While large language models have been applied to energy systems as code generators and parameter extractors, no existing implementation deploys LLMs as autonomous coordinators managing the complete workflow from natural language input to multi-appliance scheduling. This paper presents an agentic AI HEMS where LLMs autonomously coordinate multi-appliance scheduling from natural language requests to device control, achieving optimal scheduling without example demonstrations. A hierarchical architecture combining one orchestrator with three specialist agents uses the ReAct pattern for iterative reasoning, enabling dynamic coordination without hardcoded workflows while integrating Google Calendar for context-aware deadline extraction. Evaluation across three open-source models using real Austrian day-ahead electricity prices reveals substantial capability differences. Llama-3.3-70B successfully coordinates all appliances across all scenarios to match cost-optimal benchmarks computed via mixed-integer linear programming, while other models achieve perfect single-appliance performance but struggle to coordinate all appliances simultaneously. Progressive prompt engineering experiments demonstrate that analytical query handling without explicit guidance remains unreliable despite models' general reasoning capabilities. We open-source the complete system including orchestration logic, agent prompts, tools, and web interfaces to enable reproducibility, extension, and future research.
  •  

Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world

arXiv:2510.25819v1 Announce Type: cross Abstract: The rapid rise of AI agents presents urgent challenges in authentication, authorization, and identity management. Current agent-centric protocols (like MCP) highlight the demand for clarified best practices in authentication and authorization. Looking ahead, ambitions for highly autonomous agents raise complex long-term questions regarding scalable access control, agent-centric identities, AI workload differentiation, and delegated authority. This OpenID Foundation whitepaper is for stakeholders at the intersection of AI agents and access management. It outlines the resources already available for securing today's agents and presents a strategic agenda to address the foundational authentication, authorization, and identity problems pivotal for tomorrow's widespread autonomous systems.
  •  

Multi-Agent Reinforcement Learning for Market Making: Competition without Collusion

arXiv:2510.25929v1 Announce Type: cross Abstract: Algorithmic collusion has emerged as a central question in AI: Will the interaction between different AI agents deployed in markets lead to collusion? More generally, understanding how emergent behavior, be it a cartel or market dominance from more advanced bots, affects the market overall is an important research question. We propose a hierarchical multi-agent reinforcement learning framework to study algorithmic collusion in market making. The framework includes a self-interested market maker (Agent~A), which is trained in an uncertain environment shaped by an adversary, and three bottom-layer competitors: the self-interested Agent~B1 (whose objective is to maximize its own PnL), the competitive Agent~B2 (whose objective is to minimize the PnL of its opponent), and the hybrid Agent~B$^\star$, which can modulate between the behavior of the other two. To analyze how these agents shape the behavior of each other and affect market outcomes, we propose interaction-level metrics that quantify behavioral asymmetry and system-level dynamics, while providing signals potentially indicative of emergent interaction patterns. Experimental results show that Agent~B2 secures dominant performance in a zero-sum setting against B1, aggressively capturing order flow while tightening average spreads, thus improving market execution efficiency. In contrast, Agent~B$^\star$ exhibits a self-interested inclination when co-existing with other profit-seeking agents, securing dominant market share through adaptive quoting, yet exerting a milder adverse impact on the rewards of Agents~A and B1 compared to B2. These findings suggest that adaptive incentive control supports more sustainable strategic co-existence in heterogeneous agent environments and offers a structured lens for evaluating behavioral design in algorithmic trading systems.
  •  

The Quest for Reliable Metrics of Responsible AI

arXiv:2510.26007v1 Announce Type: cross Abstract: The development of Artificial Intelligence (AI), including AI in Science (AIS), should be done following the principles of responsible AI. Progress in responsible AI is often quantified through evaluation metrics, yet there has been less work on assessing the robustness and reliability of the metrics themselves. We reflect on prior work that examines the robustness of fairness metrics for recommender systems as a type of AI application and summarise their key takeaways into a set of non-exhaustive guidelines for developing reliable metrics of responsible AI. Our guidelines apply to a broad spectrum of AI applications, including AIS.
  •  
❌