❌

Reading view

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

arXiv:2605.23989v1 Announce Type: new Abstract: Agentic AI systems -- Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions -- can execute complex tasks autonomously, but their multi-step trajectories introduce new failure modes that challenge trustworthiness. This survey provides a focused examination of trustworthy agentic AI through two core dimensions that are critical for high-risk deployments: Safety and Robustness, and Privacy and System Security. For each dimension, we clarify key concepts, identify where risks emerge along the agent workflow, and summarize stage-targeted mitigation strategies. Other trustworthiness aspects (value alignment, transparency, fairness, and accountability) are discussed as relevant context rather than parallel chapters. To support consistent comparison and deployment decisions, we consolidate evaluation into a unified metrics-and-benchmarks hub, emphasizing both outcome and process signals (e.g., constraint violations, trace completeness, and adversarial success rates) and offering scenario-to-metric guidance for release gating. We conclude by outlining open challenges such as self-evolving agents, runtime monitoring and verification, privacy-preserving personalization, and the trust-utility trade-off, and present a case study of real-world security failures in open-source agentic systems. Our goal is to serve as a practical reference for researchers and practitioners building trustworthy agentic systems in high-stakes environments.
  •  

RAAP: Retrieval-Augmented Affordance Prediction with Cross-Image Action Alignment

arXiv:2603.29419v1 Announce Type: cross Abstract: Understanding object affordances is essential for enabling robots to perform purposeful and fine-grained interactions in diverse and unstructured environments. However, existing approaches either rely on retrieval, which is fragile due to sparsity and coverage gaps, or on large-scale models, which frequently mislocalize contact points and mispredict post-contact actions when applied to unseen categories, thereby hindering robust generalization. We introduce Retrieval-Augmented Affordance Prediction (RAAP), a framework that unifies affordance retrieval with alignment-based learning. By decoupling static contact localization and dynamic action direction, RAAP transfers contact points via dense correspondence and predicts action directions through a retrieval-augmented alignment model that consolidates multiple references with dual-weighted attention. Trained on compact subsets of DROID and HOI4D with as few as tens of samples per task, RAAP achieves consistent performance across unseen objects and categories, and enables zero-shot robotic manipulation in both simulation and the real world. Project website: https://github.com/SEU-VIPGroup/RAAP.
  •  

CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering

arXiv:2509.21035v2 Announce Type: replace Abstract: Knowledge graphs provide structured context for multi-hop question answering, but deployed systems must balance answer accuracy with strict latency and cost targets while preserving provenance. Static k-hop expansions and "think-longer" prompting often over-retrieve, inflate context, and yield unpredictable runtime. We introduce CLAUSE, an agentic three-agent neuro-symbolic framework that treats context construction as a sequential decision process over knowledge graphs, deciding what to expand, which paths to follow or backtrack, what evidence to keep, and when to stop. Latency (interaction steps) and prompt cost (selected tokens) are exposed as user-specified budgets or prices, allowing per-query adaptation to trade-offs among accuracy, latency, and cost without retraining. CLAUSE employs the proposed Lagrangian-Constrained Multi-Agent Proximal Policy Optimization (LC-MAPPO) algorithm to coordinate three agents: Subgraph Architect, Path Navigator, and Context Curator, so that subgraph construction, reasoning-path discovery, and evidence selection are jointly optimized under per-query resource budgets on edge edits, interaction steps, and selected tokens. Across HotpotQA, MetaQA, and FactKG, CLAUSE yields higher EM@1 while reducing subgraph growth and end-to-end latency at equal or lower token budgets. On MetaQA-2-hop, relative to the strongest RAG baseline (GraphRAG), CLAUSE achieves +39.3 EM@1 with 18.6% lower latency and 40.9% lower edge growth. The resulting contexts are compact, provenance-preserving, and deliver predictable performance under deployment constraints.
  •  

Enhancing AI-Based Tropical Cyclone Track and Intensity Forecasting via Systematic Bias Correction

arXiv:2603.22314v1 Announce Type: cross Abstract: Tropical cyclones (TCs) pose severe threats to life, infrastructure, and economies in tropical and subtropical regions, underscoring the critical need for accurate and timely forecasts of both track and intensity. Recent advances in AI-based weather forecasting have shown promise in improving TC track forecasts. However, these systems are typically trained on coarse-resolution reanalysis data (e.g., ERA5 at 0.25 degree), which constrains predicted TC positions to a fixed grid and introduces significant discretization errors. Moreover, intensity forecasting remains limited especially for strong TCs by the smoothing effect of coarse meteorological fields and the use of regression losses that bias predictions toward conditional means. To address these limitations, we propose BaguanCyclone, a novel, unified framework that integrates two key innovations: (1) a probabilistic center refinement module that models the continuous spatial distribution of TC centers, enabling finer track precision; and (2) a region-aware intensity forecasting module that leverages high-resolution internal representations within dynamically defined sub-grid zones around the TC core to better capture localized extremes. Evaluated on the global IBTrACS dataset across six major TC basins, our system consistently outperforms both operational numerical weather prediction (NWP) models and most AI-based baselines, delivering a substantial enhancement in forecast accuracy. Remarkably, BaguanCyclone excels in navigating meteorological complexities, consistently delivering accurate forecasts for re-intensification, sweeping arcs, twin cyclones, and meandering events. Our code is available at https://github.com/DAMO-DI-ML/Baguan-cyclone.
  •  

Functional-based multi-omics early prediction of radiation pneumonitis in NSCLC using AI-generated perfusion and ventilation from planning CT

Phys Med Biol. 2026 Mar 13. doi: 10.1088/1361-6560/ae5209. Online ahead of print.

ABSTRACT

ObjectiveThis study aims to develop a functional-based multi-omics model for early prediction of radiation pneumonitis (RP) by extracting radiomic and dosiomic features from functionally defined lung regions, using generated perfusion (Q) and ventilation (V) from pre-radiotherapy planning computed tomography (CT).
ApproachWe retrospectively analyzed data from 121 patients with locally advanced non-small cell lung cancer (NSCLC) treated with curative-intent IMRT between 2015 and 2019, including pre-treatment CT and dose maps. Q and V maps were generated from CT with deep learning-based and supervoxel-based approaches, respectively. Regions of interest (ROIs) combined the planning target volume (PTV) with each of three functional lung regions-high functional lung (HFL), low functional lung (LFL), and whole lung (WL)-defined by thresholds on Q and V maps. Radiomic and dosiomic features were extracted from CT and dose distributions within each ROI. For each ROI, For each ROI, three methods-radiomics (R), dosiomics (D), and dual-omics (RD)-were constructed. 13 machine learning algorithms were trained and evaluated using 10-fold cross-validation, and model performance was assessed by the average area under the receiver operating characteristic curve (AUC), accuracy, precision, recall, and F1 score. RP was defined as CTCAE grade ≥ 2.
Main resultsOf the 35 selected features, 20 were from HFL. In dual-omics models, using HFL features improved predictive performance for RP (AUC 0.879±0.105) compared to WL (AUC 0.778 ± 0.100). In HFL, the RD method outperformed both R (AUC 0.786± 0.076) and D (AUC 0.791 ± 0.107) methods. Decision curve analysis showed the dual-omics model based on HFL provided the highest net benefit across threshold probabilities.
SignificanceThis study is the first to systematically demonstrate that features extracted from CT-derived HFL capture important functional differences and provide strong predictive value for RP. Compared to conventional methods, integrating radiomics, dosiomics, and CT-based functional information further improves predictive performance.&#xD.

PMID:41825133 | DOI:10.1088/1361-6560/ae5209

  •  

Functional-based multi-omics early prediction of radiation pneumonitis in NSCLC using AI-generated perfusion and ventilation from planning CT

Phys Med Biol. 2026 Mar 13. doi: 10.1088/1361-6560/ae5209. Online ahead of print.

ABSTRACT

ObjectiveThis study aims to develop a functional-based multi-omics model for early prediction of radiation pneumonitis (RP) by extracting radiomic and dosiomic features from functionally defined lung regions, using generated perfusion (Q) and ventilation (V) from pre-radiotherapy planning computed tomography (CT).
ApproachWe retrospectively analyzed data from 121 patients with locally advanced non-small cell lung cancer (NSCLC) treated with curative-intent IMRT between 2015 and 2019, including pre-treatment CT and dose maps. Q and V maps were generated from CT with deep learning-based and supervoxel-based approaches, respectively. Regions of interest (ROIs) combined the planning target volume (PTV) with each of three functional lung regions-high functional lung (HFL), low functional lung (LFL), and whole lung (WL)-defined by thresholds on Q and V maps. Radiomic and dosiomic features were extracted from CT and dose distributions within each ROI. For each ROI, For each ROI, three methods-radiomics (R), dosiomics (D), and dual-omics (RD)-were constructed. 13 machine learning algorithms were trained and evaluated using 10-fold cross-validation, and model performance was assessed by the average area under the receiver operating characteristic curve (AUC), accuracy, precision, recall, and F1 score. RP was defined as CTCAE grade ≥ 2.
Main resultsOf the 35 selected features, 20 were from HFL. In dual-omics models, using HFL features improved predictive performance for RP (AUC 0.879±0.105) compared to WL (AUC 0.778 ± 0.100). In HFL, the RD method outperformed both R (AUC 0.786± 0.076) and D (AUC 0.791 ± 0.107) methods. Decision curve analysis showed the dual-omics model based on HFL provided the highest net benefit across threshold probabilities.
SignificanceThis study is the first to systematically demonstrate that features extracted from CT-derived HFL capture important functional differences and provide strong predictive value for RP. Compared to conventional methods, integrating radiomics, dosiomics, and CT-based functional information further improves predictive performance.&#xD.

PMID:41825133 | DOI:10.1088/1361-6560/ae5209

  •  

Facile induction of immune tolerance by an interleukin-2–TGFβ surrogate agonist

Nature, Published online: 11 March 2026; doi:10.1038/s41586-026-10208-0

A fusion protein designed to comprise IL-2 and a helminth-derived TGFβ mimic activates IL-2 and TGFβ signalling pathways in IL-2 receptor-expressing T cells and induces stable antigen-specific regulatory T cells in peripheral lymphoid organs.
  •  

Lysophosphatidylcholine acyltransferase 1 promotes head and neck squamous cell carcinoma progression by enhancing COX17-dependent oxidative phosphorylation

Cell Death Discovery, Published online: 06 March 2026; doi:10.1038/s41420-026-02994-3

Lysophosphatidylcholine acyltransferase 1 promotes head and neck squamous cell carcinoma progression by enhancing COX17-dependent oxidative phosphorylation
  •  

xLLM Technical Report

arXiv:2510.14686v2 Announce Type: replace-cross Abstract: We introduce xLLM, an intelligent and efficient Large Language Model (LLM) inference framework designed for high-performance, large-scale enterprise-grade serving, with deep optimizations for diverse AI accelerators. To address these challenges, xLLM builds a novel decoupled service-engine architecture. At the service layer, xLLM-Service features an intelligent scheduling module that efficiently processes multimodal requests and co-locates online and offline tasks through unified elastic scheduling to maximize cluster utilization. This module also relies on a workload-adaptive dynamic Prefill-Decode (PD) disaggregation policy and a novel Encode-Prefill-Decode (EPD) disaggregation policy designed for multimodal inputs. Furthermore, it incorporates a distributed architecture to provide global KV Cache management and robust fault-tolerant capabilities for high availability. At the engine layer, xLLM-Engine co-optimizes system and algorithm designs to fully saturate computing resources. This is achieved through comprehensive multi-layer execution pipeline optimizations, an adaptive graph mode and an xTensor memory management. xLLM-Engine also further integrates algorithmic enhancements such as optimized speculative decoding and dynamic EPLB, collectively serving to substantially boost throughput and inference efficiency. Extensive evaluations demonstrate that xLLM delivers significantly superior performance and resource efficiency. Under identical TPOT constraints, xLLM achieves throughput up to 1.7x that of MindIE and 2.2x that of vLLM-Ascend with Qwen-series models, while maintaining an average throughput of 1.7x that of MindIE with Deepseek-series models. xLLM framework is publicly available at https://github.com/jd-opensource/xllm and https://github.com/jd-opensource/xllm-service.
  •  

CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool Use

arXiv:2602.12268v2 Announce Type: replace Abstract: AI agents are increasingly used to solve real-world tasks by reasoning over multi-turn user interactions and invoking external tools. However, applying reinforcement learning to such settings remains difficult: realistic objectives often lack verifiable rewards and instead emphasize open-ended behaviors; moreover, RL for multi-turn, multi-step agentic tool use is still underexplored; and building and maintaining executable tool environments is costly, limiting scale and coverage. We propose CM2, an RL framework that replaces verifiable outcome rewards with checklist rewards. CM2 decomposes each turn's intended behavior into fine-grained binary criteria with explicit evidence grounding and structured metadata, turning open-ended judging into more stable classification-style decisions. To balance stability and informativeness, our method adopts a strategy of sparse reward assignment but dense evaluation criteria. Training is performed in a scalable LLM-simulated tool environment, avoiding heavy engineering for large tool sets. Experiments show that CM2 consistently improves over supervised fine-tuning. Starting from an 8B Base model and training on an 8k-example RL dataset, CM2 improves over the SFT counterpart by 8 points on tau^-Bench, by 10 points on BFCL-V4, and by 12 points on ToolSandbox. The results match or even outperform similarly sized open-source baselines, including the judging model. CM2 thus provides a scalable recipe for optimizing multi-turn, multi-step tool-using agents without relying on verifiable rewards. Code provided by the open-source community: https://github.com/namezhenzhang/CM2-RLCR-Tool-Agent.
  •  
❌