❌

Reading view

Can AI Agents Detect and Repair Artifact Drift in Network Experiments?

arXiv:2609.09849v1 Announce Type: cross Abstract: In recent years, AI agents have evolved into capable assistants that carry out multi-step tasks in digital environments. The network systems community is beginning to explore these capabilities in operational and experimental settings. However, an agent operating in network systems should not be judged solely by whether it completes the immediate task. The experiment record it modifies must also remain trustworthy. We call this property artifact integrity: the record's claims must remain supported by the available evidence, confined to the scope established by that evidence, and traceable through the artifacts that encode their support. To make this property measurable, we introduce NetArtifactBench, which tests whether AI agents can repair inconsistent records derived from public network-system artifacts while preserving claims that remain supported. The benchmark contains 52 instances with injected inconsistencies ranging from direct contradictions to unstated relations spread across several artifacts. We evaluate 23 agent configurations across three general-purpose AI agent runtimes using deterministic scoring. The average contract pass rate is 65.3 % across 5,980 outputs, but no agent runtime exceeds 30 % when repair requires recovering implicit relations and propagating changes across artifacts. These results reveal a sharp boundary between local correction and complete record-level repair. Therefore, we argue that artifact integrity should become a first-class design and evaluation requirement for AI agents operating on network systems.
  •  

Can AI Agents Deliver Verifiable Network-Wide Outcomes Across Authority Boundaries?

arXiv:2609.10181v1 Announce Type: cross Abstract: AI agents are increasingly involved in network automation, where they can initiate configuration changes through mediated operational interfaces and assess the resulting state. Nonetheless, operational networks usually span many devices and administrative domains. Realizing an operator's intent requires coordinating agents with distinct authority scopes that define the resources they can access, the operations they can invoke, and the network state they can observe. This division limits the blast radius of an erroneous action but fragments the evidence needed to assess the network-wide outcome. Successful execution of a configuration action proposed by one agent does not establish that remote devices responded as intended or that routing changes reached the required devices. A valid observation may also become stale after a subsequent change. Before the coordinated operation can be declared complete, a trusted assurance layer must collect current observations from the required scopes and determine whether they collectively support the operator's intended network-wide outcome. To address the completion admission problem, we present EvidenceNet, a runtime assurance layer for deciding whether coordinated agent operations have achieved an operator's network intent. Its broker collects the post-change observations required by a completion contract, and its admission gate checks that the evidence comes from the required scopes, remains current, and satisfies the task rules. A verifier agent provides an additional assessment of the observation content. Experiments on live routing networks show that post-change state checks recognize successful outcomes that configuration-action records alone cannot establish. Controlled interventions further show that EvidenceNet rejects completion when otherwise satisfactory observations have the wrong source, have been substituted, or are stale.
  •  

A bivalent molecular glue linking lysine acetyltransferases to oncogene-induced cell death

Chemically induced proximity of lysine acetyltransferases (KATs) with BCL6 reprograms epigenetic signaling to eliminate lymphoma tumors. Structural and mechanistic studies demonstrate that fortuitous protein-protein contacts convert proximity induction into targeted changes in chromatin, revealing a key mechanism by which small molecules can co-opt oncogenic transcriptional regulators to elicit malignant cell death.
  •  

High-salt diet in macrophage-associated metabolic disorders: Mechanisms and therapeutic implications

Chin Med J (Engl). 2026 May 19. doi: 10.1097/CM9.0000000000004098. Online ahead of print.

ABSTRACT

High-salt diet (HSD) has emerged as a prevalent environmental factor that exacerbates chronic inflammation and insulin resistance in obesity-associated type 2 diabetes (T2D) by modulating macrophage polarization, metabolic reprogramming, and epigenetic imprinting. Current evidence demonstrates that HSD activates p38/mitogen-activated protein kinase (MAPK), nuclear factor kappa-B (NF-κB), and NOD-like receptor family pyrin domain containing 3 (NLRP3) inflammasome signaling pathways, by which it drives macrophage polarization toward a proinflammatory M1 phenotype while inducing a glycolysis-dominant metabolic shift, thereby establishing a persistent "metabolic memory". Moreover, HSD orchestrates metabolic memory in macrophages through coordinated epigenetic machinery, including histone modifications (Trimethylation of histone H3 at lysine 4 [H3K4me3] and Acetylation of histone H3 at lysine 27 [H3K27ac]), DNA methylation, and noncoding RNAs (e.g., long non-coding RNA MALAT1 and miR-155), leading to sustained inflammatory phenotypes. In multiple metabolic organs (e.g., adipose tissue, liver, pancreas, and gut), the HSD-macrophage axis aggravates systemic insulin resistance through shared proinflammatory signaling and other tissue-specific mechanisms. Most importantly, therapeutic strategies targeting the NLRP3 inflammasome, metabolic pathways, and epigenetic alterations offer novel approaches for managing metabolic inflammation. Future investigations are encouraged to leverage lineage tracing, single-cell sequencing, and spatial multi-omics technologies to advance the development of precision medicine for macrophage-associated metabolic disorders.

PMID:42156155 | DOI:10.1097/CM9.0000000000004098

  •  

Pulmonary-Intestinal Axis: Shared Genetic Basis and Mediating Factors Identified Through Multi-Omics Analysis

Int J Chron Obstruct Pulmon Dis. 2026 Apr 7;21:561645. doi: 10.2147/COPD.S561645. eCollection 2026.

ABSTRACT

BACKGROUND: Chronic obstructive pulmonary disease (COPD) is a systemic condition with comorbidities beyond the lung (eg, cardiovascular and metabolic disorders), and gastrointestinal (GI) disorders are also common. The shared genetic basis of COPD-GI comorbidity and its mediating factors remain unclear. We hypothesized that COPD and GI diseases share pleiotropic genetic architecture implicating lipid-metabolic pathways, with smoking mediating part of the association.

METHODS: We analyzed publicly available European-ancestry GWAS summary statistics for COPD (Global Biobank Meta-analysis Initiative), 15 GI diseases (FinnGen), and smoking phenotypes (UK Biobank). Genetic correlation was estimated using linkage disequilibrium score regression (LDSC) and high-definition likelihood (HDL). Multi-trait analysis of GWAS (MTAG) boosted COPD discovery by leveraging genetically correlated GI traits. We integrated locus-to-gene mapping with multi-tissue expression quantitative trait loci (eQTL) and plasma protein quantitative trait loci (pQTL) evidence to prioritize shared loci, genes, and proteins. Bidirectional two-sample Mendelian randomization (MR) tested causal directions, and two-step mediation MR evaluated smoking.

RESULTS: COPD showed significant genetic correlation with nine GI diseases. We identified six comorbidity-associated loci (three with CADD > 12.37) and 13 unique candidate pleiotropic genes; APOE was supported by proteomic evidence. Enrichment analyses highlighted lipid-metabolism pathways. MR suggested COPD increases risk of gastroesophageal reflux disease (GERD), irritable bowel syndrome (IBS), acute appendicitis, and gastric ulcer, while diverticular disease showed reverse causality toward COPD. Smoking partially mediated the COPD effect on GERD, acute appendicitis, and gastric ulcer.

CONCLUSION: COPD and multiple GI disorders share a distributed pleiotropic genetic basis within the broader systemic comorbidity spectrum of COPD. Multi-omics evidence supports a genomic pulmonary-intestinal axis in which lipid metabolism and smoking-related mechanisms contribute to COPD and GI comorbidity, providing targets for risk stratification and potential intervention.

PMID:41978582 | PMC:PMC13070119 | DOI:10.2147/COPD.S561645

  •  

WIMLE: Uncertainty-Aware World Models with IMLE for Sample-Efficient Continuous Control

arXiv:2602.14351v2 Announce Type: replace-cross Abstract: Model-based reinforcement learning promises strong sample efficiency but often underperforms in practice due to compounding model error, unimodal world models that average over multi-modal dynamics, and overconfident predictions that bias learning. We introduce WIMLE, a model-based method that extends Implicit Maximum Likelihood Estimation (IMLE) to the model-based RL framework to learn stochastic, multi-modal world models without iterative sampling and to estimate predictive uncertainty via ensembles and latent sampling. During training, WIMLE weights each synthetic transition by its predicted confidence, preserving useful model rollouts while attenuating bias from uncertain predictions and enabling stable learning. Across $40$ continuous-control tasks spanning DeepMind Control, MyoSuite, and HumanoidBench, WIMLE achieves superior sample efficiency and competitive or better asymptotic performance than strong model-free and model-based baselines. Notably, on the challenging Humanoid-run task, WIMLE improves sample efficiency by over $50$\% relative to the strongest competitor, and on HumanoidBench it solves $8$ of $14$ tasks (versus $4$ for BRO and $5$ for SimbaV2). These results highlight the value of IMLE-based multi-modality and uncertainty-aware weighting for stable model-based RL.
  •  

ReFlow: Self-correction Motion Learning for Dynamic Scene Reconstruction

arXiv:2604.01561v1 Announce Type: cross Abstract: We present ReFlow, a unified framework for monocular dynamic scene reconstruction that learns 3D motion in a novel self-correction manner from raw video. Existing methods often suffer from incomplete scene initialization for dynamic regions, leading to unstable reconstruction and motion estimation, which often resorts to external dense motion guidance such as pre-computed optical flow to further stabilize and constrain the reconstruction of dynamic components. However, this introduces additional complexity and potential error propagation. To address these issues, ReFlow integrates a Complete Canonical Space Construction module for enhanced initialization of both static and dynamic regions, and a Separation-Based Dynamic Scene Modeling module that decouples static and dynamic components for targeted motion supervision. The core of ReFlow is a novel self-correction flow matching mechanism, consisting of Full Flow Matching to align 3D scene flow with time-varying 2D observations, and Camera Flow Matching to enforce multi-view consistency for static objects. Together, these modules enable robust and accurate dynamic scene reconstruction. Extensive experiments across diverse scenarios demonstrate that ReFlow achieves superior reconstruction quality and robustness, establishing a novel self-correction paradigm for monocular 4D reconstruction.
  •  

MindCube: Spatial Mental Modeling from Limited Views

arXiv:2506.21458v2 Announce Type: replace Abstract: Can Vision-Language Models (VLMs) imagine the full scene from just a few views, like humans do? Humans form spatial mental models naturally, internal representations of unseen space, to reason about layout, perspective, and motion. Our MindCube benchmark with 21,154 questions across 3,268 images exposes this critical gap, where existing VLMs exhibit near-random performance. Using MindCube, we systematically evaluate how well VLMs build robust spatial mental models through representing positions (cognitive mapping), orientations (perspective-taking), and dynamics (mental simulation for "what-if" movements). We then explore three approaches to help approximate spatial mental models in VLMs, focusing on incorporating unseen intermediate views, natural language reasoning chains, and cognitive maps. The significant improvement comes from a synergistic approach, "map-then-reason", that jointly trains the model to first generate a cognitive map and then reason upon it. By training models to reason over these internal maps, we boosted accuracy from 37.8% to 57.8% (+20.0%). Adding reinforcement learning pushed performance even further to 61.3% (+23.5%). Our key insight is that such scaffolding of spatial mental models, actively constructing and utilizing internal structured spatial representations with flexible reasoning processes, significantly improves understanding of unobservable space.
  •  

Robust transcriptomic hallmarks targeting intratumor heterogeneity in intrahepatic cholangiocarcinoma

Cell Rep Med. 2026 Mar 30:102708. doi: 10.1016/j.xcrm.2026.102708. Online ahead of print.

ABSTRACT

Intratumor heterogeneity (ITH) undermines transcriptome-based stratification in intrahepatic cholangiocarcinoma (iCCA). Here, we integrate multi-omics data from multi-region, single-region, and single-cell RNA sequencing cohorts to systematically characterize gene expression ITH. We uncover that immune and stromal heterogeneity are primary drivers of ITH, leading to misclassification of a median 27.8% of tumors by existing subtyping systems. To overcome this, we identify a low-intratumor-heterogeneity/high-intertumor-variability (LIHV) gene set and develop an ITH-insensitive classification system defining five subgroups: inflammatory (SI), metabolic (SII), atypical (SIII-1), immune-silent (SIII-2), and neurodegenerative (SIII-3). These subgroups exhibit distinct clinical outcomes, molecular features, immune landscapes, and therapeutic vulnerabilities. GPRC5A and VTCN1 serve as robust immunohistochemical biomarkers for SI and SIII tumors, while serum CEA and CA19-9 identify inflammatory iCCA. Therapeutically, HSP90 inhibition synergizes with anti-PD1 in inflammatory iCCA, whereas combined anti-PD1 and anti-TIM3 suppresses neurodegenerative iCCA. Collectively, our study provides a robust molecular framework and actionable therapeutic strategies for iCCA.

PMID:41916296 | DOI:10.1016/j.xcrm.2026.102708

  •  

Single-cell multiomics uncovers an endothelial mechanosensitive PIEZO1-IL-33 axis driving pulmonary fibrosis

Nat Commun. 2026 Mar 20;17(1):2655. doi: 10.1038/s41467-026-70193-w.

ABSTRACT

Pulmonary fibrosis represents a progressive interstitial lung disease marked by excessive extracellular matrix deposition and architectural distortion. Vascular endothelial cells critically contribute to fibrogenesis through paracrine secretion of pro-fibrotic mediators, yet their mechanobiological regulation remains elusive. Using integrated single-cell multi-omics profiling of human pulmonary fibrosis specimens and experimental fibrosis models induced by bleomycin or silica, we identify mechanosensitive Piezo1 upregulation in Endothelial cells as a hallmark of fibrotic progression. Endothelial-specific Piezo1 knockout significantly attenuates Bleomycin-induced fibrotic remodeling in male mice, establishing its pathogenic necessity. Mechanistically, PIEZO1 activation promotes pulmonary fibrosis development via CAPN2-mediated STAT3 phosphorylation, which may regulate the secretion of the pro-fibrotic molecule interleukin-33. These findings suggest that the endothelial PIEZO1-CAPN2-STAT3-IL33 axis is a potential therapeutic target for PF intervention.

PMID:41862476 | PMC:PMC13004862 | DOI:10.1038/s41467-026-70193-w

  •  

Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion

arXiv:2603.03485v1 Announce Type: cross Abstract: Recent video diffusion models have achieved impressive capabilities as large-scale generative world models. However, these models often struggle with fine-grained physical consistency, exhibiting physically implausible dynamics over time. In this work, we present \textbf{Phys4D}, a pipeline for learning physics-consistent 4D world representations from video diffusion models. Phys4D adopts \textbf{a three-stage training paradigm} that progressively lifts appearance-driven video diffusion models into physics-consistent 4D world representations. We first bootstrap robust geometry and motion representations through large-scale pseudo-supervised pretraining, establishing a foundation for 4D scene modeling. We then perform physics-grounded supervised fine-tuning using simulation-generated data, enforcing temporally consistent 4D dynamics. Finally, we apply simulation-grounded reinforcement learning to correct residual physical violations that are difficult to capture through explicit supervision. To evaluate fine-grained physical consistency beyond appearance-based metrics, we introduce a set of \textbf{4D world consistency evaluation} that probe geometric coherence, motion stability, and long-horizon physical plausibility. Experimental results demonstrate that Phys4D substantially improves fine-grained spatiotemporal and physical consistency compared to appearance-driven baselines, while maintaining strong generative performance. Our project page is available at https://sensational-brioche-7657e7.netlify.app/
  •  

Learning to Generate and Extract: A Multi-Agent Collaboration Framework For Zero-shot Document-level Event Arguments Extraction

arXiv:2603.02909v2 Announce Type: replace-cross Abstract: Document-level event argument extraction (DEAE) is essential for knowledge acquisition, aiming to extract participants of events from documents . In the zero-shot setting, existing methods employ LLMs to generate synthetic data to address the challenge posed by the scarcity of annotated data. However, relying solely on Event-type-only prompts makes it difficult for the generated content to accurately capture the contextual and structural relationships of unseen events. Moreover, ensuring the reliability and usability of synthetic data remains a significant challenge due to the absence of quality evaluation mechanisms. To this end, we introduce a multi-agent collaboration framework for zero-shot document-level event argument extraction (ZS-DEAE), which simulates the human collaborative cognitive process of "Propose-Evaluate-Revise." Specifically, the framework comprises a generation agent and an evaluation agent. The generation agent synthesizes data for unseen events by leveraging knowledge from seen events, while the evaluation agent extracts arguments from the synthetic data and assesses their semantic consistency with the context. The evaluation results are subsequently converted into reward signals, with event structure constraints incorporated into the reward design to enable iterative optimization of both agents via reinforcement learning.In three zero-shot scenarios constructed from the RAMS and WikiEvents datasets, our method achieves improvements both in data generation quality and argument extraction performance, while the generated data also effectively enhances the zero-shot performance of other DEAE models.
  •  

Learning to Generate and Extract: A Multi-Agent Collaboration Framework For Zero-shot Document-level Event Arguments Extraction

arXiv:2603.02909v1 Announce Type: cross Abstract: Document-level event argument extraction (DEAE) is essential for knowledge acquisition, aiming to extract participants of events from documents.In the zero-shot setting, existing methods employ LLMs to generate synthetic data to address the challenge posed by the scarcity of annotated data. However, relying solely on Event-type-only prompts makes it difficult for the generated content to accurately capture the contextual and structural relationships of unseen events. Moreover, ensuring the reliability and usability of synthetic data remains a significant challenge due to the absence of quality evaluation mechanisms. To this end, we introduce a multi-agent collaboration framework for zero-shot document-level event argument extraction (ZS-DEAE), which simulates the human collaborative cognitive process of "Propose-Evaluate-Revise." Specifically, the framework comprises a generation agent and an evaluation agent. The generation agent synthesizes data for unseen events by leveraging knowledge from seen events, while the evaluation agent extracts arguments from the synthetic data and assesses their semantic consistency with the context. The evaluation results are subsequently converted into reward signals, with event structure constraints incorporated into the reward design to enable iterative optimization of both agents via reinforcement learning.In three zero-shot scenarios constructed from the RAMS and WikiEvents datasets, our method achieves improvements both in data generation quality and argument extraction performance, while the generated data also effectively enhances the zero-shot performance of other DEAE models.
  •  

MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language Models

arXiv:2509.23725v3 Announce Type: replace Abstract: Answering complex medical questions requires not only domain expertise and patient-specific information, but also structured and multi-perspective reasoning. Existing multi-agent approaches often rely on fixed roles or shallow interaction prompts, limiting their ability to detect and resolve fine-grained logical inconsistencies. To address this, we propose \textsc{MedLA}, a logic-driven multi-agent framework built on large language models. Each agent organizes its reasoning process into an explicit logical tree based on syllogistic triads (major premise, minor premise, and conclusion), enabling transparent inference and premise-level alignment. Agents engage in a multi-round, graph-guided discussion to compare and iteratively refine their logic trees, achieving consensus through error correction and contradiction resolution. We demonstrate that \textsc{MedLA} consistently outperforms both static role-based systems and single-agent baselines on challenging benchmarks such as MedDDx and standard medical QA tasks. Furthermore, \textsc{MedLA} scales effectively across both open-source and commercial LLM backbones, achieving state-of-the-art performance and offering a generalizable paradigm for trustworthy medical reasoning.
  •  

EBPO: Empirical Bayes Shrinkage for Stabilizing Group-Relative Policy Optimization

arXiv:2602.05165v3 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for enhancing the reasoning capabilities of Large Language Models (LLMs). However, dominant approaches like Group Relative Policy Optimization (GRPO) face critical stability challenges: they suffer from high estimator variance under computational constraints (small group sizes) and vanishing gradient signals in saturated failure regimes where all responses yield identical zero rewards. To address this, we propose Empirical Bayes Policy Optimization (EBPO), a novel framework that regularizes local group-based baselines by borrowing strength from the policy's accumulated global statistics. Instead of estimating baselines in isolation, EBPO employs a shrinkage estimator that dynamically balances local group statistics with a global prior updated via Welford's online algorithm. Theoretically, we demonstrate that EBPO guarantees strictly lower Mean Squared Error (MSE), bounded entropy decay, and non-vanishing penalty signals in failure scenarios compared to GRPO. Empirically, EBPO consistently outperforms GRPO and other established baselines across diverse benchmarks, including AIME and OlympiadBench. Notably, EBPO exhibits superior training stability, achieving high-performance gains even with small group sizes, and benefits significantly from difficulty-stratified curriculum learning.
  •  

WIMLE: Uncertainty-Aware World Models with IMLE for Sample-Efficient Continuous Control

arXiv:2602.14351v1 Announce Type: cross Abstract: Model-based reinforcement learning promises strong sample efficiency but often underperforms in practice due to compounding model error, unimodal world models that average over multi-modal dynamics, and overconfident predictions that bias learning. We introduce WIMLE, a model-based method that extends Implicit Maximum Likelihood Estimation (IMLE) to the model-based RL framework to learn stochastic, multi-modal world models without iterative sampling and to estimate predictive uncertainty via ensembles and latent sampling. During training, WIMLE weights each synthetic transition by its predicted confidence, preserving useful model rollouts while attenuating bias from uncertain predictions and enabling stable learning. Across $40$ continuous-control tasks spanning DeepMind Control, MyoSuite, and HumanoidBench, WIMLE achieves superior sample efficiency and competitive or better asymptotic performance than strong model-free and model-based baselines. Notably, on the challenging Humanoid-run task, WIMLE improves sample efficiency by over $50$\% relative to the strongest competitor, and on HumanoidBench it solves $8$ of $14$ tasks (versus $4$ for BRO and $5$ for SimbaV2). These results highlight the value of IMLE-based multi-modality and uncertainty-aware weighting for stable model-based RL.
  •  
❌