❌

Reading view

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

arXiv:2605.23989v1 Announce Type: new Abstract: Agentic AI systems -- Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions -- can execute complex tasks autonomously, but their multi-step trajectories introduce new failure modes that challenge trustworthiness. This survey provides a focused examination of trustworthy agentic AI through two core dimensions that are critical for high-risk deployments: Safety and Robustness, and Privacy and System Security. For each dimension, we clarify key concepts, identify where risks emerge along the agent workflow, and summarize stage-targeted mitigation strategies. Other trustworthiness aspects (value alignment, transparency, fairness, and accountability) are discussed as relevant context rather than parallel chapters. To support consistent comparison and deployment decisions, we consolidate evaluation into a unified metrics-and-benchmarks hub, emphasizing both outcome and process signals (e.g., constraint violations, trace completeness, and adversarial success rates) and offering scenario-to-metric guidance for release gating. We conclude by outlining open challenges such as self-evolving agents, runtime monitoring and verification, privacy-preserving personalization, and the trust-utility trade-off, and present a case study of real-world security failures in open-source agentic systems. Our goal is to serve as a practical reference for researchers and practitioners building trustworthy agentic systems in high-stakes environments.
  •  

LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition

arXiv:2605.24005v1 Announce Type: new Abstract: The evolution of Large Language Model (LLM) reasoning is bottlenecked by the scarcity of high-quality process data. While self-alignment via endogenous rewards offers a solution, mining valid supervision faces three challenges: (1) Label Noise via Mimetic Bias, where rewards prioritize statistical likelihood over logical truth, creating a "correctness illusion" that masks compounding errors; (2) Coarse-Grained Supervision, where sparse global outcomes (e.g., in GRPO) fail to provide granular guidance, treating reasoning chains as monolithic; and (3) Distributional Collapse, where signals fail to generalize without amplifying pre-training biases. To address these, we introduce LC-ERD (Logic-Consistent Endogenous Reward Decomposition), a framework framing self-alignment as latent structure mining. We derive a Variational Logic Potential by aggregating consensus from the model's Latent Logic Expertise (LLE) to denoise the reasoning manifold, and introduce a Multi-Agent Value Decomposition protocol based on the IGM principle to quantify individual step utility. Experiments show LC-ERD delivers a robust self-evolution path, uncovering trade-offs between logic consistency and accuracy while identifying high-value reasoning patterns missed by standard rewards. Our code is available at https://github.com/Reinhardmannn/LC-ERD.
  •  

SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking

arXiv:2605.25160v1 Announce Type: new Abstract: Mobile GUI agents powered by large language models have progressed rapidly, creating urgent needs for realistic and comprehensive evaluation. Existing benchmarks prioritize reproducibility but are often limited to open-source apps or file-operation tasks for the difficulty of constructing rewards on real applications, leaving a gap between benchmark settings and real-world usage. Moreover, most benchmarks focus on basic grounding and navigation, with limited coverage of complex, long-horizon interactions. To address these limitations, we introduce SimuWoB, a fully synthetic benchmark for mobile GUI agents with 120 challenging tasks spanning diverse types and difficulty levels. We build a robust virtual environment generation framework that synthesizes high-fidelity tasks and environments, and automatically provides valid rewards for each task. Each environment is deployed as a backend-free webpage accessible via URL, enabling efficient and reproducible evaluation. We conduct comprehensive experiments on several state-of-the-art mobile GUI agents. The average success rate is only 27.92%, dropping to 17.82% on long-horizon tasks, which reveals substantial weaknesses in current agents under complex scenarios. Evaluation result comparison with real-world sample tasks demonstrate that agent assessments based on our synthetic environment generalize well. We further provide diagnostic insights across key capability dimensions and discuss implications for future mobile GUI agent development.
  •  

MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research

arXiv:2605.26114v1 Announce Type: new Abstract: We present MobileGym, a browser-hosted, lightweight, fully controllable environment for everyday mobile use, targeting interaction fidelity without replicating proprietary backends. It enables two capabilities previously out of reach for everyday apps: verifiable outcome signals through deterministic state-based judging over structured JSON state, and scalable online RL through low-cost parallel rollouts. The full environment state is captured, configured, forked, and compared as structured JSON, and a single server can host hundreds of parallel instances, with about 400 MB memory per instance and about 3 s cold start. A layered state model and a declarative task-definition framework keep state programmability and task creation practical at scale, and a single programmatic judging mechanism delivers both deterministic evaluation verdicts and dense RL rewards. The accompanying MobileGym-Bench provides 416 parameterized task templates, including 256 test and 160 train templates, over 28 apps, with deterministic judges and a structured AnswerSheet protocol that avoids free-text matching failures. In a Sim-to-Real case study, GRPO on Qwen3-VL-4B-Instruct gains +12.8 percentage points on the 256-task test set, and on a 59-task real-device signal subset, real-device execution retains 95.1% of the simulation-side training gain. Project page: https://mobilegym.github.io.
  •  

Unlocking the Potential of Continual Model Merging: An ODE Perspective

arXiv:2605.19409v3 Announce Type: replace-cross Abstract: Continual Model Merging (CMM) enables rapid customization of foundation models by sequentially incorporating task-adapted models without repeated retraining. However, existing merging rules usually update the deployed model through fixed algebraic or projection-based operations, providing limited control over how much previously accumulated knowledge should be retained relative to the incoming task model. This limitation leads to unstable retention and performance degradation in long task streams, and becomes more pronounced when tasks have heterogeneous utilities. We propose ODE-driven Merging (ODE-M), a controllable framework that formulates each continual merge as a trajectory in parameter space rather than a one-step endpoint update. Motivated by mode connectivity, ODE-M constructs a barrier-aware trajectory using a rectified time-dependent velocity field, where lightweight first-order feedback from a small calibration set suppresses loss-increasing motion while preserving progress toward the incoming model. The next merged model is then obtained by selecting an operating point along this trajectory through a utility-aware time schedule, providing an explicit mechanism for balancing retained historical knowledge and incoming task expertise. Extensive experiments on standard CMM benchmarks show that ODE-M consistently improves over strong continual merging baselines across CLIP ViT backbones, stream lengths, and heterogeneous task-utility settings.
  •  

AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild

arXiv:2605.22715v2 Announce Type: replace-cross Abstract: As wearable and mobile devices become increasingly embedded in daily life, they offer a practical way to continuously sense human motion in the wild. But inertial signals are highly dependent on the sensing setup, including body location, mounting position, sensor orientation, device hardware, and sampling protocol. This setup dependence makes it difficult to learn motion representations that transfer across devices and datasets, and limits the broader use of wearable IMUs beyond closed-set recognition. We introduce AnyMo, a geometry-aware framework for setup-agnostic human motion modeling. AnyMo uses physics-grounded IMU simulation over dense body-surface placements to generate diverse and plausible synthetic signals, pre-trains a graph encoder from paired synthetic placement views and masked partial observations, tokenizes multi-position IMU into full-body motion tokens, and aligns these tokens with an LLM for motion-language understanding. We evaluate AnyMo on three complementary tasks: zero-shot activity recognition across 14 unseen downstream datasets, cross-modal retrieval, and wearable IMU motion captioning, where it improves average Accuracy/F1/R@2 by 11.7\%/11.6\%/22.6\% on HAR, increases zero-shot IMU-to-text and text-to-IMU retrieval MRR by 15.9\% and 28.6\%, respectively, and improves zero-shot captioning BERT-F1 by 18.8\%. These results support AnyMo as a generalist model for wearable motion understanding in the wild. Project page: https://baiyuchen.com/project/AnyMo.
  •  

Molecular basis for methylation-sensitive editing by Cas9

Nature, Published online: 15 April 2026; doi:10.1038/s41586-026-10384-z

ThermoCas9, a genome-editing enzyme that is sensitive to the DNA methylation status of the target locus, is characterized and shows promise for targeting hypomethylated DNA regions in cancer cells.
  •  

EBV strain interacts with host HLA to drive nasopharyngeal carcinoma risk

Nature, Published online: 15 April 2026; doi:10.1038/s41586-026-10416-8

A genome-to-genome association study identifies host and viral risk factors that interact to drive nasopharyngeal carcinoma endemicity in southern China.
  •  

Transplantation of encapsulated mitochondria alleviates dysfunction in mitochondrial and Parkinson’s disease models

A mitochondrial transplantation approach rescues mitochondrial deficiency and prevents mitochondrial DNA depletion syndrome, Leigh syndrome, and Parkinson’s disease in cellular and mouse models.
  •  

Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models

arXiv:2604.03302v1 Announce Type: cross Abstract: While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focus. Despite their improvements, current MLLMs struggle significantly with high-level physics reasoning. In this work, we investigate the first step of physical reasoning, i.e., intuitive physics understanding, revealing substantial limitations in understanding the dynamics of continuum objects. To isolate and evaluate this specific capability, we introduce two fundamental benchmark tasks: Next Frame Selection (NFS) and Temporal Coherence Verification (TCV). Our experiments demonstrate that even state-of-the-art MLLMs perform poorly on these foundational tasks. To address this limitation, we propose Scene Dynamic Field (SDF), a concise approach that leverages physics simulators within a multi-task fine-tuning framework. SDF substantially improves performance, achieving up to 20.7% gains on fluid tasks while showing strong generalization to unseen physical domains. This work not only highlights a critical gap in current MLLMs but also presents a promising cost-efficient approach for developing more physically grounded MLLMs. Our code and data are available at https://github.com/andylinx/Scene-Dynamic-Field.
  •  

DP-OPD: Differentially Private On-Policy Distillation for Language Models

arXiv:2604.04461v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly adapted to proprietary and domain-specific corpora that contain sensitive information, creating a tension between formal privacy guarantees and efficient deployment through model compression. Differential privacy (DP), typically enforced via DP-SGD, provides record-level protection but often incurs substantial utility loss in autoregressive generation, where optimization noise can amplify exposure bias and compounding errors along long rollouts. Existing approaches to private distillation either apply DP-SGD to both teacher and student, worsening computation and the privacy--utility tradeoff, or rely on DP synthetic text generation from a DP-trained teacher, avoiding DP on the student at the cost of DP-optimizing a large teacher and introducing an offline generation pipeline. We propose \textbf{Differentially Private On-Policy Distillation (DP-OPD)}, a synthesis-free framework that enforces privacy solely through DP-SGD on the student while leveraging a frozen teacher to provide dense token-level targets on \emph{student-generated} trajectories. DP-OPD instantiates this idea via \emph{private generalized knowledge distillation} on continuation tokens. Under a strict privacy budget ($\varepsilon=2.0$), DP-OPD improves perplexity over DP fine-tuning and off-policy DP distillation, and outperforms synthesis-based DP distillation (Yelp: 44.15$\rightarrow$41.68; BigPatent: 32.43$\rightarrow$30.63), while substantially simplifying the training pipeline. In particular, \textbf{DP-OPD collapses private compression into a single DP student-training loop} by eliminating DP teacher training and offline synthetic text generation. Code will be released upon publication at https://github.com/khademfatemeh/dp_opd.
  •  

Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation

arXiv:2601.00263v2 Announce Type: replace-cross Abstract: Counterfactuals refer to minimally edited inputs that cause a model's prediction to change, serving as a promising approach to explaining the model's behavior. Large language models (LLMs) excel at generating English counterfactuals and demonstrate multilingual proficiency. However, their effectiveness in generating multilingual counterfactuals remains unclear. To this end, we conduct a comprehensive study on multilingual counterfactuals. We first conduct automatic evaluations on both directly generated counterfactuals in the target languages and those derived via English translation across six languages. Although translation-based counterfactuals offer higher validity than their directly generated counterparts, they demand substantially more modifications and still fall short of matching the quality of the original English counterfactuals. Second, we find the patterns of edits applied to high-resource European-language counterfactuals to be remarkably similar, suggesting that cross-lingual perturbations follow common strategic principles. Third, we identify and categorize four main types of errors that consistently appear in the generated counterfactuals across languages. Finally, we reveal that multilingual counterfactual data augmentation (CDA) yields larger model performance improvements than cross-lingual CDA, especially for lower-resource languages. Yet, the imperfections of the generated counterfactuals limit gains in model performance and robustness.
  •  

Decoding macrophage heterogeneity in the pulmonary fibrosis lung cancer transition

Front Immunol. 2026 Mar 20;17:1787094. doi: 10.3389/fimmu.2026.1787094. eCollection 2026.

ABSTRACT

Pulmonary fibrosis (PF) significantly increases the risk of lung cancer (LC), but the mechanisms underlying this transition remain unclear. This overview positions macrophage heterogeneity as a central node within the PF-LC continuum. First, we describe important subpopulations of profibrotic and pro-tumor macrophages, including SPP1+, MERTK+, TREM2+, and MARCO+ cells, using high-resolution spatial and single-cell omics technologies. Next, we analyze the fundamental mechanisms that determine their function: the fibrotic microenvironment (e.g., extracellular matrix stiffness, hypoxia) induces profound metabolic reprogramming (e.g., Warburg effect, lipid peroxidation) and stabilizes epigenetic memory (e.g., DNA methylation, histone modifications), locking them into a pathogenic state. This reprogramming occurs through two main pathways: (1) metabolic reprogramming, characterized by aerobic glycolytic conversion and dysregulated lipid metabolism, which stimulates both pathogenic functions and suppression of T cell activity; (2) Epigenetic modifications, including stabilized alterations in DNA methylation, histone modifications, and superactivator patterns, which maintain cells in a tumor-promoting phenotype. As central nodes of communication, these macrophages interact pathologically with fibroblasts and epithelial cells through secreted factors and extracellular vesicles, forming self-reinforcing feedback loops that promote disease progression. We are studying the crucial role of new technologies, particularly multi-omic spatial models and high-precision organoids, in fostering mechanistic discoveries. These discoveries pave the way for new macrophage-focused therapeutic strategies, including the precise stratification of patients using biomarkers from liquid biopsies (such as soluble SPP1 and MARCO) and the development of targeted drug delivery systems for the selective modulation of macrophage function, thus establishing a new paradigm for therapeutic interventions in pulmonary fibrosis with concomitant lung cancer.

PMID:41939908 | PMC:PMC13046558 | DOI:10.3389/fimmu.2026.1787094

  •  

Decoding macrophage heterogeneity in the pulmonary fibrosis lung cancer transition

Front Immunol. 2026 Mar 20;17:1787094. doi: 10.3389/fimmu.2026.1787094. eCollection 2026.

ABSTRACT

Pulmonary fibrosis (PF) significantly increases the risk of lung cancer (LC), but the mechanisms underlying this transition remain unclear. This overview positions macrophage heterogeneity as a central node within the PF-LC continuum. First, we describe important subpopulations of profibrotic and pro-tumor macrophages, including SPP1+, MERTK+, TREM2+, and MARCO+ cells, using high-resolution spatial and single-cell omics technologies. Next, we analyze the fundamental mechanisms that determine their function: the fibrotic microenvironment (e.g., extracellular matrix stiffness, hypoxia) induces profound metabolic reprogramming (e.g., Warburg effect, lipid peroxidation) and stabilizes epigenetic memory (e.g., DNA methylation, histone modifications), locking them into a pathogenic state. This reprogramming occurs through two main pathways: (1) metabolic reprogramming, characterized by aerobic glycolytic conversion and dysregulated lipid metabolism, which stimulates both pathogenic functions and suppression of T cell activity; (2) Epigenetic modifications, including stabilized alterations in DNA methylation, histone modifications, and superactivator patterns, which maintain cells in a tumor-promoting phenotype. As central nodes of communication, these macrophages interact pathologically with fibroblasts and epithelial cells through secreted factors and extracellular vesicles, forming self-reinforcing feedback loops that promote disease progression. We are studying the crucial role of new technologies, particularly multi-omic spatial models and high-precision organoids, in fostering mechanistic discoveries. These discoveries pave the way for new macrophage-focused therapeutic strategies, including the precise stratification of patients using biomarkers from liquid biopsies (such as soluble SPP1 and MARCO) and the development of targeted drug delivery systems for the selective modulation of macrophage function, thus establishing a new paradigm for therapeutic interventions in pulmonary fibrosis with concomitant lung cancer.

PMID:41939908 | PMC:PMC13046558 | DOI:10.3389/fimmu.2026.1787094

  •  

Restoring circadian rhythms in the hypothalamic paraventricular nucleus reverses aging biomarkers and extends lifespan in male mice

Enhancing circadian amplitude in mouse hypothalamic paraventricular nucleus neurons by 3′-deoxyadenosine treatment alleviates age-related pathologies and extends lifespan.
  •  

Multi-Omics Characterization of Lactate-Associated Molecular Subtypes in Lung Cancer Suggests a Role for DKK1 in Lactate-Linked Migration, Invasion, and Lactylation Programs

Cancers (Basel). 2026 Feb 25;18(5):735. doi: 10.3390/cancers18050735.

ABSTRACT

BACKGROUND: Lactate accumulation is increasingly recognized as a feature of tumor metabolic reprogramming that can coincide with immune dysregulation and aggressive phenotypes. The prognostic and immunologic relevance of lactate-associated heterogeneity in lung cancer remains to be clarified.

METHODS: We curated lactate-related genes and identified prognostic candidates in lung cancer cohorts. Consensus clustering was applied to define lactate-associated molecular subtypes, followed by characterization of survival and tumor microenvironment features. A LASSO-based gene signature was developed to generate an individual-level risk score and an integrated nomogram. Multi-omics analyses were used to evaluate concordance between transcriptomic and proteomic alterations. Single-cell transcriptomic data were analyzed to explore cellular heterogeneity in lactate-related programs. In vitro assays evaluated the response of candidate genes to lactate exposure and assessed cell migration and invasion under proliferation-inhibited conditions after genetic perturbation.

RESULTS: Two lactate-associated molecular subtypes were identified with distinct overall survival and divergent immune microenvironment features. Subtype 1 was associated with better outcomes and a more immune-inflamed profile, whereas Subtype 2 was associated with poorer outcomes and a myeloid-enriched, immunosuppressive contexture. Pathway analyses indicated subtype-associated differences in extracellular matrix-related processes and apoptosis-associated signaling. We developed an 11-gene prognostic signature and nomogram that stratified patients by risk across TCGA and GEO cohorts. Multi-omics integration highlighted ANLN, FGA, and DKK1 as consistently dysregulated at both transcript and protein levels. Among these candidates, DKK1 showed lactate-responsive induction in vitro. DKK1 perturbation altered lactate-enhanced migratory and invasive phenotypes and was accompanied by changes in intracellular lactate levels and global protein lactylation, supporting a potential feedforward relationship between lactate exposure, DKK1 expression, and lactylation.

CONCLUSIONS: This study characterizes lactate-associated molecular heterogeneity in lung cancer and provides a lactate-related subtype framework and prognostic risk model for patient stratification. The findings nominate DKK1 as a lactate-responsive candidate linked to migration/invasion phenotypes and lactate/lactylation changes in vitro.

PMID:41827671 | PMC:PMC12985219 | DOI:10.3390/cancers18050735

  •  

Multi-Omics Characterization of Lactate-Associated Molecular Subtypes in Lung Cancer Suggests a Role for DKK1 in Lactate-Linked Migration, Invasion, and Lactylation Programs

Cancers (Basel). 2026 Feb 25;18(5):735. doi: 10.3390/cancers18050735.

ABSTRACT

BACKGROUND: Lactate accumulation is increasingly recognized as a feature of tumor metabolic reprogramming that can coincide with immune dysregulation and aggressive phenotypes. The prognostic and immunologic relevance of lactate-associated heterogeneity in lung cancer remains to be clarified.

METHODS: We curated lactate-related genes and identified prognostic candidates in lung cancer cohorts. Consensus clustering was applied to define lactate-associated molecular subtypes, followed by characterization of survival and tumor microenvironment features. A LASSO-based gene signature was developed to generate an individual-level risk score and an integrated nomogram. Multi-omics analyses were used to evaluate concordance between transcriptomic and proteomic alterations. Single-cell transcriptomic data were analyzed to explore cellular heterogeneity in lactate-related programs. In vitro assays evaluated the response of candidate genes to lactate exposure and assessed cell migration and invasion under proliferation-inhibited conditions after genetic perturbation.

RESULTS: Two lactate-associated molecular subtypes were identified with distinct overall survival and divergent immune microenvironment features. Subtype 1 was associated with better outcomes and a more immune-inflamed profile, whereas Subtype 2 was associated with poorer outcomes and a myeloid-enriched, immunosuppressive contexture. Pathway analyses indicated subtype-associated differences in extracellular matrix-related processes and apoptosis-associated signaling. We developed an 11-gene prognostic signature and nomogram that stratified patients by risk across TCGA and GEO cohorts. Multi-omics integration highlighted ANLN, FGA, and DKK1 as consistently dysregulated at both transcript and protein levels. Among these candidates, DKK1 showed lactate-responsive induction in vitro. DKK1 perturbation altered lactate-enhanced migratory and invasive phenotypes and was accompanied by changes in intracellular lactate levels and global protein lactylation, supporting a potential feedforward relationship between lactate exposure, DKK1 expression, and lactylation.

CONCLUSIONS: This study characterizes lactate-associated molecular heterogeneity in lung cancer and provides a lactate-related subtype framework and prognostic risk model for patient stratification. The findings nominate DKK1 as a lactate-responsive candidate linked to migration/invasion phenotypes and lactate/lactylation changes in vitro.

PMID:41827671 | PMC:PMC12985219 | DOI:10.3390/cancers18050735

  •  

Multi-omics investigation of benzo[a]pyrene in gastric cancer: comprehensive network toxicology, machine learning and molecular docking approaches

Mol Divers. 2026 Mar 12. doi: 10.1007/s11030-026-11508-3. Online ahead of print.

ABSTRACT

Gastric cancer (GC) risk is shaped by environmental exposures such as benzo[a]pyrene (BaP). Here, we systematically identified BaP-toxicological targets and dissected their contribution to GC development. BaP-related targets were independently predicted with stringent filters from ChEMBL, Similarity Ensemble Approach (SEA) and PharmMapper databases, while GC-related targets were mined from the Comparative Toxicogenomics Database (CTD), GeneCards and OMIM databases. Overlapping targets were subjected to protein-protein interaction (PPI) network construction, functional enrichment analysis and molecular docking. We then integrated multi-omics data using ten clustering algorithms to identify the consensus GC subtypes, which were subsequently employed 101 machine learning combinations to develop a consensus benzo[a]pyrene-related signature (CBRS) for GC patients. As a result, we identified seven hub toxicological targets: ALB, HSP90AA1, ESR1, INS, TP53, TNF, and EGFR, underscoring their potential central roles in BaP-driven GC pathogenesis. These targets are enriched in the MAPK, Lipid and atherosclerosis, and PI3K-Akt signaling pathway. The BaP-toxicological classifiers and the CBRS prognostic model could provide useful support for risk stratification and inform personalized therapeutic strategies for GC patients. Molecular docking results suggest that BaP exhibits relatively strong binding affinity with these key toxicological targets, potentially implicating their involvement in BaP-induced gastric cancer toxicity. Therefore, this study integrates multi-dimensional omics data with advanced machine learning algorithms to establish a comprehensive analytical framework for the toxicological effects of between BaP and GC, which transcends the limitations of traditional analyses and offers unprecedented insights and evidence chains for elucidating the pathogenesis of GC.

PMID:41817952 | DOI:10.1007/s11030-026-11508-3

  •  
❌