❌

Normal view

The Denario project: Deep knowledge AI agents for scientific discovery

arXiv:2510.26887v1 Announce Type: new Abstract: We present Denario, an AI multi-agent system designed to serve as a scientific research assistant. Denario can perform many different tasks, such as generating ideas, checking the literature, developing research plans, writing and executing code, making plots, and drafting and reviewing a scientific paper. The system has a modular architecture, allowing it to handle specific tasks, such as generating an idea, or carrying out end-to-end scientific analysis using Cmbagent as a deep-research backend. In this work, we describe in detail Denario and its modules, and illustrate its capabilities by presenting multiple AI-generated papers generated by it in many different scientific disciplines such as astrophysics, biology, biophysics, biomedical informatics, chemistry, material science, mathematical physics, medicine, neuroscience and planetary science. Denario also excels at combining ideas from different disciplines, and we illustrate this by showing a paper that applies methods from quantum physics and machine learning to astrophysical data. We report the evaluations performed on these papers by domain experts, who provided both numerical scores and review-like feedback. We then highlight the strengths, weaknesses, and limitations of the current system. Finally, we discuss the ethical implications of AI-driven research and reflect on how such technology relates to the philosophy of science. We publicly release the code at https://github.com/AstroPilot-AI/Denario. A Denario demo can also be run directly on the web at https://huggingface.co/spaces/astropilot-ai/Denario, and the full app will be deployed on the cloud.

Glia: A Human-Inspired AI for Automated Systems Design and Optimization

arXiv:2510.27176v1 Announce Type: new Abstract: Can an AI autonomously design mechanisms for computer systems on par with the creativity and reasoning of human experts? We present Glia, an AI architecture for networked systems design that uses large language models (LLMs) in a human-inspired, multi-agent workflow. Each agent specializes in reasoning, experimentation, and analysis, collaborating through an evaluation framework that grounds abstract reasoning in empirical feedback. Unlike prior ML-for-systems methods that optimize black-box policies, Glia generates interpretable designs and exposes its reasoning process. When applied to a distributed GPU cluster for LLM inference, it produces new algorithms for request routing, scheduling, and auto-scaling that perform at human-expert levels in significantly less time, while yielding novel insights into workload behavior. Our results suggest that by combining reasoning LLMs with structured experimentation, an AI can produce creative and understandable designs for complex systems problems.

ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use

arXiv:2510.27363v1 Announce Type: new Abstract: Recently, large language models (LLMs) have demonstrated remarkable problem-solving capabilities by autonomously integrating with external tools for collaborative reasoning. However, due to the inherently complex and diverse nature of multimodal information, enabling multimodal large language models (MLLMs) to flexibly and efficiently utilize external tools during reasoning remains an underexplored challenge. In this work, we introduce ToolScope, an agentic framework designed to unify global planning with local multimodal perception, adopting a specialized Perceive tool to mitigates visual context degradation in long-horizon VQA task. ToolScope comprises three primary components: the Global Navigator, the Agentic Executor, and the Response Synthesizer. The Global Navigator functions as a "telescope", offering high-level strategic guidance. The Agentic Executor operates iteratively to augment MLLM with local perception through the integration of external tools-Search, Code, and Perceive. Finally, the Response Synthesizer consolidates and organizes the reasoning process into a coherent, user-friendly output. We evaluate ToolScope on four VQA benchmarks across diverse domains, including VQA 2.0, ScienceQA, MAT-Search and MathVista. It demonstrates strong generalization capabilities, achieving an average performance improvement of up to +6.69% across all datasets.

VeriMoA: A Mixture-of-Agents Framework for Spec-to-HDL Generation

arXiv:2510.27617v1 Announce Type: new Abstract: Automation of Register Transfer Level (RTL) design can help developers meet increasing computational demands. Large Language Models (LLMs) show promise for Hardware Description Language (HDL) generation, but face challenges due to limited parametric knowledge and domain-specific constraints. While prompt engineering and fine-tuning have limitations in knowledge coverage and training costs, multi-agent architectures offer a training-free paradigm to enhance reasoning through collaborative generation. However, current multi-agent approaches suffer from two critical deficiencies: susceptibility to noise propagation and constrained reasoning space exploration. We propose VeriMoA, a training-free mixture-of-agents (MoA) framework with two synergistic innovations. First, a quality-guided caching mechanism to maintain all intermediate HDL outputs and enables quality-based ranking and selection across the entire generation process, encouraging knowledge accumulation over layers of reasoning. Second, a multi-path generation strategy that leverages C++ and Python as intermediate representations, decomposing specification-to-HDL translation into two-stage processes that exploit LLM fluency in high-resource languages while promoting solution diversity. Comprehensive experiments on VerilogEval 2.0 and RTLLM 2.0 benchmarks demonstrate that VeriMoA achieves 15--30% improvements in Pass@1 across diverse LLM backbones, especially enabling smaller models to match larger models and fine-tuned alternatives without requiring costly training.

Frame Semantic Patterns for Identifying Underreporting of Notifiable Events in Healthcare: The Case of Gender-Based Violence

arXiv:2510.26969v1 Announce Type: cross Abstract: We introduce a methodology for the identification of notifiable events in the domain of healthcare. The methodology harnesses semantic frames to define fine-grained patterns and search them in unstructured data, namely, open-text fields in e-medical records. We apply the methodology to the problem of underreporting of gender-based violence (GBV) in e-medical records produced during patients' visits to primary care units. A total of eight patterns are defined and searched on a corpus of 21 million sentences in Brazilian Portuguese extracted from e-SUS APS. The results are manually evaluated by linguists and the precision of each pattern measured. Our findings reveal that the methodology effectively identifies reports of violence with a precision of 0.726, confirming its robustness. Designed as a transparent, efficient, low-carbon, and language-agnostic pipeline, the approach can be easily adapted to other health surveillance contexts, contributing to the broader, ethical, and explainable use of NLP in public health systems.

Detecting Data Contamination in LLMs via In-Context Learning

arXiv:2510.27055v1 Announce Type: cross Abstract: We present Contamination Detection via Context (CoDeC), a practical and accurate method to detect and quantify training data contamination in large language models. CoDeC distinguishes between data memorized during training and data outside the training distribution by measuring how in-context learning affects model performance. We find that in-context examples typically boost confidence for unseen datasets but may reduce it when the dataset was part of training, due to disrupted memorization patterns. Experiments show that CoDeC produces interpretable contamination scores that clearly separate seen and unseen datasets, and reveals strong evidence of memorization in open-weight models with undisclosed training corpora. The method is simple, automated, and both model- and dataset-agnostic, making it easy to integrate with benchmark evaluations.

MARIA: A Framework for Marginal Risk Assessment without Ground Truth in AI Systems

arXiv:2510.27163v1 Announce Type: cross Abstract: Before deploying an AI system to replace an existing process, it must be compared with the incumbent to ensure improvement without added risk. Traditional evaluation relies on ground truth for both systems, but this is often unavailable due to delayed or unknowable outcomes, high costs, or incomplete data, especially for long-standing systems deemed safe by convention. The more practical solution is not to compute absolute risk but the difference between systems. We therefore propose a marginal risk assessment framework, that avoids dependence on ground truth or absolute risk. It emphasizes three kinds of relative evaluation methodology, including predictability, capability and interaction dominance. By shifting focus from absolute to relative evaluation, our approach equips software teams with actionable guidance: identifying where AI enhances outcomes, where it introduces new risks, and how to adopt such systems responsibly.

MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models

arXiv:2510.27196v1 Announce Type: cross Abstract: The proliferation of memes on social media necessitates the capabilities of multimodal Large Language Models (mLLMs) to effectively understand multimodal harmfulness. Existing evaluation approaches predominantly focus on mLLMs' detection accuracy for binary classification tasks, which often fail to reflect the in-depth interpretive nuance of harmfulness across diverse contexts. In this paper, we propose MemeArena, an agent-based arena-style evaluation framework that provides a context-aware and unbiased assessment for mLLMs' understanding of multimodal harmfulness. Specifically, MemeArena simulates diverse interpretive contexts to formulate evaluation tasks that elicit perspective-specific analyses from mLLMs. By integrating varied viewpoints and reaching consensus among evaluators, it enables fair and unbiased comparisons of mLLMs' abilities to interpret multimodal harmfulness. Extensive experiments demonstrate that our framework effectively reduces the evaluation biases of judge agents, with judgment results closely aligning with human preferences, offering valuable insights into reliable and comprehensive mLLM evaluations in multimodal harmfulness understanding. Our code and data are publicly available at https://github.com/Lbotirx/MemeArena.

Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments

arXiv:2510.27287v1 Announce Type: cross Abstract: Enterprise systems are crucial for enhancing productivity and decision-making among employees and customers. Integrating LLM based systems into enterprise systems enables intelligent automation, personalized experiences, and efficient information retrieval, driving operational efficiency and strategic growth. However, developing and evaluating such systems is challenging due to the inherent complexity of enterprise environments, where data is fragmented across multiple sources and governed by sophisticated access controls. We present EnterpriseBench, a comprehensive benchmark that simulates enterprise settings, featuring 500 diverse tasks across software engineering, HR, finance, and administrative domains. Our benchmark uniquely captures key enterprise characteristics including data source fragmentation, access control hierarchies, and cross-functional workflows. Additionally, we provide a novel data generation pipeline that creates internally consistent enterprise tasks from organizational metadata. Experiments with state-of-the-art LLM agents demonstrate that even the most capable models achieve only 41.8% task completion, highlighting significant opportunities for improvement in enterprise-focused AI systems.

Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning

arXiv:2510.27606v1 Announce Type: cross Abstract: Spatial understanding remains a weakness of Large Vision-Language Models (LVLMs). Existing supervised fine-tuning (SFT) and recent reinforcement learning with verifiable rewards (RLVR) pipelines depend on costly supervision, specialized tools, or constrained environments that limit scale. We introduce Spatial-SSRL, a self-supervised RL paradigm that derives verifiable signals directly from ordinary RGB or RGB-D images. Spatial-SSRL automatically formulates five pretext tasks that capture 2D and 3D spatial structure: shuffled patch reordering, flipped patch recognition, cropped patch inpainting, regional depth ordering, and relative 3D position prediction. These tasks provide ground-truth answers that are easy to verify and require no human or LVLM annotation. Training on our tasks substantially improves spatial reasoning while preserving general visual capabilities. On seven spatial understanding benchmarks in both image and video settings, Spatial-SSRL delivers average accuracy gains of 4.63% (3B) and 3.89% (7B) over the Qwen2.5-VL baselines. Our results show that simple, intrinsic supervision enables RLVR at scale and provides a practical route to stronger spatial intelligence in LVLMs.

Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models

arXiv:2510.27629v1 Announce Type: cross Abstract: Open-weight bio-foundation models present a dual-use dilemma. While holding great promise for accelerating scientific research and drug development, they could also enable bad actors to develop more deadly bioweapons. To mitigate the risk posed by these models, current approaches focus on filtering biohazardous data during pre-training. However, the effectiveness of such an approach remains unclear, particularly against determined actors who might fine-tune these models for malicious use. To address this gap, we propose \eval, a framework to evaluate the robustness of procedures that are intended to reduce the dual-use capabilities of bio-foundation models. \eval assesses models' virus understanding through three lenses, including sequence modeling, mutational effects prediction, and virulence prediction. Our results show that current filtering practices may not be particularly effective: Excluded knowledge can be rapidly recovered in some cases via fine-tuning, and exhibits broader generalizability in sequence modeling. Furthermore, dual-use signals may already reside in the pretrained representations, and can be elicited via simple linear probing. These findings highlight the challenges of data filtering as a standalone procedure, underscoring the need for further research into robust safety and security strategies for open-weight bio-foundation models.
  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • A Survey of AI Scientists Guiyao Tie · Pan Zhou · Lichao Sun
    arXiv:2510.23045v3 Announce Type: replace Abstract: Artificial intelligence is undergoing a profound transition from a computational instrument to an autonomous originator of scientific knowledge. This emerging paradigm, the AI scientist, is architected to emulate the complete scientific workflow-from initial hypothesis generation to the final synthesis of publishable findings-thereby promising to fundamentally reshape the pace and scale of discovery. However, the rapid and unstructured prolife
     

A Survey of AI Scientists

arXiv:2510.23045v3 Announce Type: replace Abstract: Artificial intelligence is undergoing a profound transition from a computational instrument to an autonomous originator of scientific knowledge. This emerging paradigm, the AI scientist, is architected to emulate the complete scientific workflow-from initial hypothesis generation to the final synthesis of publishable findings-thereby promising to fundamentally reshape the pace and scale of discovery. However, the rapid and unstructured proliferation of these systems has created a fragmented research landscape, obscuring overarching methodological principles and developmental trends. This survey provides a systematic and comprehensive synthesis of this domain by introducing a unified, six-stage methodological framework that deconstructs the end-to-end scientific process into: Literature Review, Idea Generation, Experimental Preparation, Experimental Execution, Scientific Writing, and Paper Generation. Through this analytical lens, we chart the field's evolution from early Foundational Modules (2022-2023) to integrated Closed-Loop Systems (2024), and finally to the current frontier of Scalability, Impact, and Human-AI Collaboration (2025-present). By rigorously synthesizing these developments, this survey not only clarifies the current state of autonomous science but also provides a critical roadmap for overcoming remaining challenges in robustness and governance, ultimately guiding the next generation of systems toward becoming trustworthy and indispensable partners in human scientific inquiry.

A Systematic Literature Review of Spatio-Temporal Graph Neural Network Models for Time Series Forecasting and Classification

arXiv:2410.22377v3 Announce Type: replace-cross Abstract: In recent years, spatio-temporal graph neural networks (GNNs) have attracted considerable interest in the field of time series analysis, due to their ability to capture, at once, dependencies among variables and across time points. The objective of this systematic literature review is hence to provide a comprehensive overview of the various modeling approaches and application domains of GNNs for time series classification and forecasting. A database search was conducted, and 366 papers were selected for a detailed examination of the current state-of-the-art in the field. This examination is intended to offer to the reader a comprehensive review of proposed models, links to related source code, available datasets, benchmark models, and fitting results. All this information is hoped to assist researchers in their studies. To the best of our knowledge, this is the first and broadest systematic literature review presenting a detailed comparison of results from current spatio-temporal GNN models applied to different domains. In its final part, this review discusses current limitations and challenges in the application of spatio-temporal GNNs, such as comparability, reproducibility, explainability, poor information capacity, and scalability. This paper is complemented by a GitHub repository at https://github.com/FlaGer99/SLR-Spatio-Temporal-GNN.git providing additional interactive tools to further explore the presented findings.

A Multi-Stage Framework with Taxonomy-Guided Reasoning for Occupation Classification Using Large Language Models

arXiv:2503.12989v3 Announce Type: replace-cross Abstract: Automatically annotating job data with standardized occupations from taxonomies, known as occupation classification, is crucial for labor market analysis. However, this task is often hindered by data scarcity and the challenges of manual annotations. While large language models (LLMs) hold promise due to their extensive world knowledge and in-context learning capabilities, their effectiveness depends on their knowledge of occupational taxonomies, which remains unclear. In this study, we assess the ability of LLMs to generate precise taxonomic entities from taxonomy, highlighting their limitations, especially for smaller models. To address these challenges, we propose a multi-stage framework consisting of inference, retrieval, and reranking stages, which integrates taxonomy-guided reasoning examples to enhance performance by aligning outputs with taxonomic knowledge. Evaluations on a large-scale dataset show that our framework not only enhances occupation and skill classification tasks, but also provides a cost-effective alternative to frontier models like GPT-4o, significantly reducing computational costs while maintaining strong performance. This makes it a practical and scalable solution for occupation classification and related tasks across LLMs.

BALR-SAM: Boundary-Aware Low-Rank Adaptation of SAM for Resource-Efficient Medical Image Segmentation

arXiv:2509.24204v2 Announce Type: replace-cross Abstract: Vision foundation models like the Segment Anything Model (SAM), pretrained on large-scale natural image datasets, often struggle in medical image segmentation due to a lack of domain-specific adaptation. In clinical practice, fine-tuning such models efficiently for medical downstream tasks with minimal resource demands, while maintaining strong performance, is challenging. To address these issues, we propose BALR-SAM, a boundary-aware low-rank adaptation framework that enhances SAM for medical imaging. It combines three tailored components: (1) a Complementary Detail Enhancement Network (CDEN) using depthwise separable convolutions and multi-scale fusion to capture boundary-sensitive features essential for accurate segmentation; (2) low-rank adapters integrated into SAM's Vision Transformer blocks to optimize feature representation and attention for medical contexts, while simultaneously significantly reducing the parameter space; and (3) a low-rank tensor attention mechanism in the mask decoder, cutting memory usage by 75% and boosting inference speed. Experiments on standard medical segmentation datasets show that BALR-SAM, without requiring prompts, outperforms several state-of-the-art (SOTA) methods, including fully fine-tuned MedSAM, while updating just 1.8% (11.7M) of its parameters.

International expert consensus on the clinical integration of circulating tumor cells in solid tumors

Eur J Cancer. 2025 Dec 9;231:116050. doi: 10.1016/j.ejca.2025.116050. Epub 2025 Oct 20.

ABSTRACT

BACKGROUND: Circulating tumor cells (CTCs) are a versatile biomarker in solid tumors. Extensive research supports their clinical relevance and led to regulatory approval in breast, prostate, and colorectal cancers. However, clinical adoption remains limited mainly due to the lack of consensus and standardized technologies. Additionally, CTC research lacks unified direction. To address these gaps, an international expert panel was established to assess the current and future clinical utility of CTCs.

METHODS: A panel of 11 CTC experts identified key areas of controversy, informing a structured survey distributed to 55 international multidisciplinary experts. Consensus was predefined as ≥ 70 % agreement. Areas without consensus were discussed in a virtual meeting, leading to final statements on the clinical integration of CTCs.

RESULTS: Thirty-seven experts completed the survey. Consensus was reached on the clinical utility of CTCs for prognosis and treatment monitoring in metastatic breast (BC) and prostate (PC) cancers, including AR-V7 testing in metastatic castration-resistant PC for therapy selection. In other tumors, CTCs remain investigational. Experts agreed that while clinical utility is not yet established in early-stage disease, CTCs show promise in early BC, especially combined with cell-free DNA (cfDNA) for minimal residual disease detection. CellSearch® is currently the only platform with high-level evidence for clinical use, though emerging technologies are promising. Key challenges include improving detection sensitivity/specificity, standardizing workflows, generating robust data, and clinician education. Experts emphasized shifting from enumeration to phenotypic and molecular characterization, particularly for treatment guidance, and highlighted the complementary role of CTCs and cfDNA, advocating for integrated liquid biopsy approaches.

CONCLUSIONS: This consensus offers practical guidance for clinical integration of CTCs and outlines strategic research priorities to unlock their full potential in precision oncology.

PMID:41172567 | DOI:10.1016/j.ejca.2025.116050

Animal models in tuberculosis metabolomics: a systematic review of current evidence and the road to translational relevance

Front Mol Biosci. 2025 Oct 15;12:1688882. doi: 10.3389/fmolb.2025.1688882. eCollection 2025.

ABSTRACT

BACKGROUND: Animal models are important for tuberculosis (TB) research, offering controlled settings to study disease mechanisms. However, their ability to replicate TB-induced metabolic responses in humans is uncertain. This systematic review evaluated the current use of animal models in metabolomics studies aimed at characterising active pulmonary TB.

METHODS: PubMed, Scopus, and Web of Science were systematically searched for metabolomics studies of pulmonary TB in humans and animal models, following PRISMA guidelines. Eligible studies were screened, and quality was assessed using QUDOMICS and STAIR tools. Data were synthesised by species, sample matrix, experimental design, and reported differential metabolites. Differential metabolite names were compared between species and subjected to pathway analysis in MetaboAnalyst 6.0.

RESULTS: Of the 80 eligible studies, nine involved animal models, predominantly mice. These models captured only 4.7% of human TB-associated differential metabolites, with the highest overlap (3.8%) in mouse lung tissue. Despite low concordance at metabolite level, conserved disruptions were observed in amino acid, glutathione, and one-carbon metabolism pathways. Interspecies variation was evident, influenced by host species, sample matrix, infection protocol, and analytical method.

CONCLUSION: Animal models partially replicated key metabolic features of human TB, particularly at the pathway level. However, variability across studies hampers current translational interpretation. Broader model use, standardised protocols, and integrated multi-platform omics approaches are needed to improve the relevance and comparability of animal models in TB metabolomics research.

PMID:41169614 | PMC:PMC12568366 | DOI:10.3389/fmolb.2025.1688882

Digital Health Technology Compliance With Clinical Safety Standards In the National Health Service in England: National Cross-Sectional Study

Background: To be authorized for use in the National Health Service (NHS) in England, digital health technologies (DHTs) must meet 2 mandatory clinical risk management standards, Data Coordination Board (DCB) 0129 and 0160, demonstrating that risks from design and use have been assessed and mitigated. NHS organizations must not procure a DHT without DCB0129 assurance and must not deploy one without DCB0160 assurance. Despite legal requirement, no public data exist on how many DHTs are in use in the NHS or how many are assured. Objective: This study aimed to determine the number of DHTs in use in the NHS in England and assess their assurance status against mandated clinical safety standards. Methods: In early 2025, 239 NHS organizations in England received a freedom of information notice requesting information on the number of DHTs they were using and their assurance against DCB0129 and DCB0160 standards. Results: Of the 239 NHS organizations, 204 (85.4%) responded, of which 178 (87.3%) provided full or partial data, covering 14,747 DHT deployments. The mean number of deployed DHTs per organization was 82.8 (SD 146.1; 95% CI 61.4-104.3) with substantial variation between NHS provider trusts (mean 107.1, SD 161.1; 95% CI 79.8-134.3), ambulance trusts (mean 13.0, SD 8.2; 95% CI 7.6-18.4), and integrated care boards (mean 8.1, SD 16.0; 95% CI 2.8-13.5). Overall organizational compliance rates were low, with a median of 25.6% (IQR 7.8%-55.7%) deployed DHTs being fully assured; for NHS provider trusts compliance was lower at 24.5% (IQR 8.1%-50%). A total of 13 (6.4%) of the 204 organizations reported that all their DHTs were fully assured, while 16 (7.8%) reported that none were assured. Across all DHTs with reported assurance data, 17.3% (95% CI 16.6%-18.1%) were fully assured against both standards, 13.3% were partially assured against one standard, and 70.1% (95% CI 69.1%-71.1%) had no documented assurance. Conclusions: This is the first study to quantify both the scale of DHT deployment in NHS organizations in England and the extent of compliance with mandatory safety standards. More than 10,000 DHTs currently in use lack documented assurance against clinical safety standards. In a typical NHS trust, 3 out of 4 digital tools influencing patient care do not demonstrate compliance with minimum legal or clinical safety requirements. These findings raise significant concerns about the risks posed to patients by these technologies; the capacity of organizations to assess and mitigate them; and the legal ramifications of when, not if, harm occurs. Crucially, failure to assure digital technologies poses a significant risk to one of the core ambitions of the NHS 10-Year Health Plan for England; safely transitioning from analogue to digital care models. These findings are unlikely to be unique to the NHS and should prompt health care systems worldwide to assess the risks posed by their DHT deployments.

An Agentic Framework for Rapid Deployment of Edge AI Solutions in Industry 5.0

arXiv:2510.25813v1 Announce Type: new Abstract: We present a novel framework for Industry 5.0 that simplifies the deployment of AI models on edge devices in various industrial settings. The design reduces latency and avoids external data transfer by enabling local inference and real-time processing. Our implementation is agent-based, which means that individual agents, whether human, algorithmic, or collaborative, are responsible for well-defined tasks, enabling flexibility and simplifying integration. Moreover, our framework supports modular integration and maintains low resource requirements. Preliminary evaluations concerning the food industry in real scenarios indicate improved deployment time and system adaptability performance. The source code is publicly available at https://github.com/AI-REDGIO-5-0/ci-component.

Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world

arXiv:2510.25819v1 Announce Type: cross Abstract: The rapid rise of AI agents presents urgent challenges in authentication, authorization, and identity management. Current agent-centric protocols (like MCP) highlight the demand for clarified best practices in authentication and authorization. Looking ahead, ambitions for highly autonomous agents raise complex long-term questions regarding scalable access control, agent-centric identities, AI workload differentiation, and delegated authority. This OpenID Foundation whitepaper is for stakeholders at the intersection of AI agents and access management. It outlines the resources already available for securing today's agents and presents a strategic agenda to address the foundational authentication, authorization, and identity problems pivotal for tomorrow's widespread autonomous systems.
❌