❌

Normal view

Optimizing Service Operations via LLM-Powered Multi-Agent Simulation

arXiv:2604.04383v1 Announce Type: new Abstract: Service system performance depends on how participants respond to design choices, but modeling these responses is hard due to the complexity of human behavior. We introduce an LLM-powered multi-agent simulation (LLM-MAS) framework for optimizing service operations. We pose the problem as stochastic optimization with decision-dependent uncertainty: design choices are embedded in prompts and shape the distribution of outcomes from interacting LLM-powered agents. By embedding key numerical information in prompts and extracting it from LLM-generated text, we model this uncertainty as a controlled Markov chain. We develop an on-trajectory learning algorithm that, on a single simulation run, simultaneously constructs zeroth-order gradient estimates and updates design parameters to optimize steady-state performance. We also incorporate variance reduction techniques. In a sustainable supply chain application, our method outperforms benchmarks, including blackbox optimization and using LLMs as numerical solvers or as role-playing system designers. A case study on optimal contest design with real behavioral data shows that LLM-MAS is both as a cost-effective evaluator of known designs and an exploratory tool that can uncover strong designs overlooked by traditional approaches.

GUIDE: Interpretable GUI Agent Evaluation via Hierarchical Diagnosis

arXiv:2604.04399v1 Announce Type: new Abstract: Evaluating GUI agents presents a distinct challenge: trajectories are long, visually grounded, and open-ended, yet evaluation must be both accurate and interpretable. Existing approaches typically apply a single holistic judgment over the entire action-observation sequence-a strategy that proves unreliable on long-horizon tasks and yields binary verdicts offering no insight into where or why an agent fails. This opacity limits the utility of evaluation as a diagnostic tool for agent development. We introduce GUIDE (GUI Understanding and Interpretable Diagnostic Evaluation), a framework that decomposes trajectory assessment into three sequential stages mirroring the compositional structure of GUI tasks. Trajectory Segmentation partitions the full trace into semantically coherent subtask units. Subtask Diagnosis evaluates each unit in context, assigning a completion verdict and generating a structured error analysis with corrective recommendations. Overall Summary aggregates per-subtask diagnoses into a task-level judgment. By operating on bounded subtask segments rather than full trajectories, GUIDE mitigates the context overload that degrades existing evaluators as task complexity grows. We validate GUIDE on three benchmarks: an industrial e-commerce dataset of 932 trajectories, AGENTREWARDBENCH spanning five web agent tasks with 1302 trajectories, and AndroidBench for mobile device control. Across all settings, GUIDE substantially outperforms existing evaluators-achieving up to 5.35 percentage points higher accuracy than the strongest baseline-while producing structured diagnostic reports that directly inform agent improvement.

Mind Your HEARTBEAT! Claw Background Execution Inherently Enables Silent Memory Pollution

arXiv:2603.23064v3 Announce Type: replace-cross Abstract: We identify a critical security vulnerability in mainstream Claw personal AI agents: untrusted content encountered during heartbeat-driven background execution can silently pollute agent memory and subsequently influence user-facing behavior without the user's awareness. This vulnerability arises from an architectural design shared across the Claw ecosystem: heartbeat background execution runs in the same session as user-facing conversation, so content ingested from any external source monitored in the background (including email, message channels, news feeds, code repositories, and social platforms) can enter the same memory context used for foreground interaction, often with limited user visibility and without clear source provenance. We formalize this process as an Exposure (E) $\rightarrow$ Memory (M) $\rightarrow$ Behavior (B) pathway: misinformation encountered during heartbeat execution enters the agent's short-term session context, potentially gets written into long-term memory, and later shapes downstream user-facing behavior. We instantiate this pathway in an agent-native social setting using MissClaw, a controlled research replica of Moltbook. We find that (1) social credibility cues, especially perceived consensus, are the dominant driver of short-term behavioral influence, with misleading rates up to 61%; (2) routine memory-saving behavior can promote short-term pollution into durable long-term memory at rates up to 91%, with cross-session behavioral influence reaching 76%; (3) under naturalistic browsing with content dilution and context pruning, pollution still crosses session boundaries. Overall, prompt injection is not required: ordinary social misinformation is sufficient to silently shape agent memory and behavior under heartbeat-driven background execution.

Lifting Unlabeled Internet-level Data for 3D Scene Understanding

arXiv:2604.01907v1 Announce Type: cross Abstract: Annotated 3D scene data is scarce and expensive to acquire, while abundant unlabeled videos are readily available on the internet. In this paper, we demonstrate that carefully designed data engines can leverage web-curated, unlabeled videos to automatically generate training data, to facilitate end-to-end models in 3D scene understanding alongside human-annotated datasets. We identify and analyze bottlenecks in automated data generation, revealing critical factors that determine the efficiency and effectiveness of learning from unlabeled data. To validate our approach across different perception granularities, we evaluate on three tasks spanning low-level perception, i.e., 3D object detection and instance segmentation, to high-evel reasoning, i.e., 3D spatial Visual Question Answering (VQA) and Vision-Lanugage Navigation (VLN). Models trained on our generated data demonstrate strong zero-shot performance and show further improvement after finetuning. This demonstrates the viability of leveraging readily available web data as a path toward more capable scene understanding systems.

DVM: A Bytecode Virtual Machine Approach for Dynamic Tensor Computation

arXiv:2603.24239v2 Announce Type: replace-cross Abstract: Dynamism is common in AI computation, e.g., the dynamic tensor shapes and the dynamic control flows in models. Due to the long compilation time, existing runtime compilation damages the model efficiency, while the offline compilers either suffer from the long compilation time and device memory footprint to cover all the possible execution instances of a dynamic model, or sacrifice optimization opportunities for usability. In this paper, we rethink the feasibility of runtime compilation for dynamic models and identify that the key for it to work is to speed up the compilation or hide the compilation overhead. To do this, we propose a real-time compiler, DVM. In DVM, we design a runtime operator compiler based on a bytecode virtual machine to perform effective and efficient compilation for each dynamic operator instance given its input. Specifically, instead of compiling programs into machine code, we encode the operator program into bytecode on the CPU and decode the bytecode into virtual instructions for direct execution on the NPU. Based on the runtime operator compiler, we further propose an operator fuser, which performs symbol-deduction-based fusion on static graphs and runtime fusion on dynamic graphs. Both pattern- and stacking-based fusion are supported to increase fusion opportunities. Evaluation on operators, subgraphs, and models shows that, compared with TorchInductor, PyTorch-eager and MindSpore-graph-O0, we are up to 11.77$\times$ better in terms of the operator/model efficiency and up to 5 orders of magnitude faster in terms of the maximum compilation time.

Let the Agent Steer: Closed-Loop Ranking Optimization via Influence Exchange

arXiv:2603.27765v2 Announce Type: replace Abstract: Recommendation ranking is fundamentally an influence allocation problem: a sorting formula distributes ranking influence among competing factors, and the business outcome depends on finding the optimal "exchange rates" among them. However, offline proxy metrics systematically misjudge how influence reallocation translates to online impact, with asymmetric bias across metrics that a single calibration factor cannot correct. We present Sortify, the first fully autonomous LLM-driven ranking optimization agent deployed in a large-scale production recommendation system. The agent reframes ranking optimization as continuous influence exchange, closing the full loop from diagnosis to parameter deployment without human intervention. It addresses structural problems through three mechanisms: (1) a dual-channel framework grounded in Savage's Subjective Expected Utility (SEU) that decouples offline-online transfer correction (Belief channel) from constraint penalty adjustment (Preference channel); (2) an LLM meta-controller operating on framework-level parameters rather than low-level search variables; (3) a persistent Memory DB with 7 relational tables for cross-round learning. Its core metric, Influence Share, provides a decomposable measure where all factor contributions sum to exactly 100%. Sortify has been deployed across two markets. In Country A, the agent pushed GMV from -3.6% to +9.2% within 7 rounds with peak orders reaching +12.5%. In Country B, a cold-start deployment achieved +4.15% GMV/UU and +3.58% Ads Revenue in a 7-day A/B test, leading to full production rollout.

InCoder-32B: Code Foundation Model for Industrial Scenarios

arXiv:2603.16790v3 Announce Type: replace-cross Abstract: Recent code large language models have achieved remarkable progress on general programming tasks. Nevertheless, their performance degrades significantly in industrial scenarios that require reasoning about hardware semantics, specialized language constructs, and strict resource constraints. To address these challenges, we introduce InCoder-32B (Industrial-Coder-32B), the first 32B-parameter code foundation model unifying code intelligence across chip design, GPU kernel optimization, embedded systems, compiler optimization, and 3D modeling. By adopting an efficient architecture, we train InCoder-32B from scratch with general code pre-training, curated industrial code annealing, mid-training that progressively extends context from 8K to 128K tokens with synthetic industrial reasoning data, and post-training with execution-grounded verification. We conduct extensive evaluation on 14 mainstream general code benchmarks and 9 industrial benchmarks spanning 4 specialized domains. Results show InCoder-32B achieves highly competitive performance on general tasks while establishing strong open-source baselines across industrial domains.

A unified deep learning framework for cross-platform harmonization of multi-tracer PET quantification in neurodegenerative disease

npj Digital Medicine, Published online: 30 March 2026; doi:10.1038/s41746-026-02570-0

A unified deep learning framework for cross-platform harmonization of multi-tracer PET quantification in neurodegenerative disease

Targeting sialic acid metabolism: a therapeutic strategy against gastric cancer driven by WZ35

Cell Oncol (Dordr). 2026 Mar 23;49(2):60. doi: 10.1007/s13402-026-01194-6.

ABSTRACT

Glycolytic reprogramming is closely associated with the occurrence and progression of gastric cancer. Specifically, the energy derived from glucose metabolism and the cellular proteins by its intermediate products influence gastric cancer development. However, as an important branch of glucose metabolism, sialic acid metabolism and its mediated sialylation modifications remain insufficiently studied in gastric cancer, and their specific relationship with malignant tumor progression requires further exploration. This study employed a multi‑omics approach, integrating metabolomics, single‑cell RNA sequencing, and bulk RNA sequencing analyses, to investigate the metabolic landscape of gastric cancer and its associated alterations. The results indicated that sialic acid is a characteristic metabolite in malignant gastric cancer tissues. It modulates biological functions such as immune response, proliferative activity, and metabolic remodeling within gastric cancer tissues by influencing sialylation modifications. Furthermore, we identified the drug WZ35, which can inhibit the malignant proliferation of gastric cancer by targeting both sialic acid metabolism and sialylated protein modifications. We put forward a conjecture that the metabolism and modification of sialic acid promote the malignant development of gastric cancer, and we discovered that the drug WZ35 has an inhibitory effect on the sialic acid metabolism of gastric cancer.

GRAPHICAL ABSTRACT:

PMID:41870836 | PMC:PMC13009457 | DOI:10.1007/s13402-026-01194-6

MERIT: Memory-Enhanced Retrieval for Interpretable Knowledge Tracing

arXiv:2603.22289v1 Announce Type: cross Abstract: Knowledge Tracing (KT) models students' evolving knowledge states to predict future performance, serving as a foundation for personalized education. While traditional deep learning models achieve high accuracy, they often lack interpretability. Large Language Models (LLMs) offer strong reasoning capabilities but struggle with limited context windows and hallucinations. Furthermore, existing LLM-based methods typically require expensive fine-tuning, limiting scalability and adaptability to new data. We propose MERIT (Memory-Enhanced Retrieval for Interpretable Knowledge Tracing), a training-free framework combining frozen LLM reasoning with structured pedagogical memory. Rather than updating parameters, MERIT transforms raw interaction logs into an interpretable memory bank. The framework uses semantic denoising to categorize students into latent cognitive schemas and constructs a paradigm bank where representative error patterns are analyzed offline to generate explicit Chain-of-Thought (CoT) rationales. During inference, a hierarchical routing mechanism retrieves relevant contexts, while a logic-augmented module applies semantic constraints to calibrate predictions. By grounding the LLM in interpretable memory, MERIT achieves state-of-the-art performance on real-world datasets without gradient updates. This approach reduces computational costs and supports dynamic knowledge updates, improving the accessibility and transparency of educational diagnosis.

UniQueR: Unified Query-based Feedforward 3D Reconstruction

arXiv:2603.22851v1 Announce Type: cross Abstract: We present UniQueR, a unified query-based feedforward framework for efficient and accurate 3D reconstruction from unposed images. Existing feedforward models such as DUSt3R, VGGT, and AnySplat typically predict per-pixel point maps or pixel-aligned Gaussians, which remain fundamentally 2.5D and limited to visible surfaces. In contrast, UniQueR formulates reconstruction as a sparse 3D query inference problem. Our model learns a compact set of 3D anchor points that act as explicit geometric queries, enabling the network to infer scene structure, including geometry in occluded regions--in a single forward pass. Each query encodes spatial and appearance priors directly in global 3D space (instead of per-frame camera space) and spawns a set of 3D Gaussians for differentiable rendering. By leveraging unified query interactions across multi-view features and a decoupled cross-attention design, UniQueR achieves strong geometric expressiveness while substantially reducing memory and computational cost. Experiments on Mip-NeRF 360 and VR-NeRF demonstrate that UniQueR surpasses state-of-the-art feedforward methods in both rendering quality and geometric accuracy, using an order of magnitude fewer primitives than dense alternatives.

Mind Your HEARTBEAT! Claw Background Execution Inherently Enables Silent Memory Pollution

arXiv:2603.23064v2 Announce Type: cross Abstract: We identify a critical security vulnerability in mainstream Claw personal AI agents: untrusted content encountered during heartbeat-driven background execution can silently pollute agent memory and subsequently influence user-facing behavior without the user's awareness. This vulnerability arises from an architectural design shared across the Claw ecosystem: heartbeat background execution runs in the same session as user-facing conversation, so content ingested from any external source monitored in the background (including email, message channels, news feeds, code repositories, and social platforms) can enter the same memory context used for foreground interaction, often with limited user visibility and without clear source provenance. We formalize this process as an Exposure (E) $\rightarrow$ Memory (M) $\rightarrow$ Behavior (B) pathway: misinformation encountered during heartbeat execution enters the agent's short-term session context, potentially gets written into long-term memory, and later shapes downstream user-facing behavior. We instantiate this pathway in an agent-native social setting using MissClaw, a controlled research replica of Moltbook. We find that (1) social credibility cues, especially perceived consensus, are the dominant driver of short-term behavioral influence, with misleading rates up to 61%; (2) routine memory-saving behavior can promote short-term pollution into durable long-term memory at rates up to 91%, with cross-session behavioral influence reaching 76%; (3) under naturalistic browsing with content dilution and context pruning, pollution still crosses session boundaries. Overall, prompt injection is not required: ordinary social misinformation is sufficient to silently shape agent memory and behavior under heartbeat-driven background execution.

Dataset Distillation-based Hybrid Federated Learning on Non-IID Data

arXiv:2409.17517v3 Announce Type: replace-cross Abstract: In federated learning, the heterogeneity of client data has a great impact on the performance of model training. Many heterogeneity issues in this process are raised by non-independently and identically distributed (non-IID) data. To address the issue of label distribution skew, we propose a hybrid federated learning framework called HFLDD, which integrates dataset distillation to generate approximately independent and equally distributed (IID) data, thereby improving the performance of model training. In particular, we partition the clients into heterogeneous clusters, where the data labels among different clients within a cluster are unbalanced while the data labels among different clusters are balanced. The cluster heads collect distilled data from the corresponding cluster members, and conduct model training in collaboration with the server. This training process is like traditional federated learning on IID data, and hence effectively alleviates the impact of non-IID data on model training. We perform a comprehensive analysis of the convergence behavior, communication overhead, and computational complexity of the proposed HFLDD. Extensive experimental results based on multiple public datasets demonstrate that when data labels are severely imbalanced, the proposed HFLDD outperforms the baseline methods in terms of both test accuracy and communication cost.

FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control

arXiv:2603.12612v1 Announce Type: cross Abstract: Scaling Maximum Entropy Reinforcement Learning (RL) to high-dimensional humanoid control remains a formidable challenge, as the ``curse of dimensionality'' induces severe exploration inefficiency and training instability in expansive action spaces. Consequently, recent high-throughput paradigms have largely converged on deterministic policy gradients combined with massive parallel simulation. We challenge this compromise with FastDSAC, a framework that effectively unlocks the potential of maximum entropy stochastic policies for complex continuous control. We introduce Dimension-wise Entropy Modulation (DEM) to dynamically redistribute the exploration budget and enforce diversity, alongside a continuous distributional critic tailored to ensure value fidelity and mitigate high-dimensional value overestimation. Extensive evaluations on HumanoidBench and other continuous control tasks demonstrate that rigorously designed stochastic policies can consistently match or outperform deterministic baselines, achieving notable gains of 180\% and 400\% on the challenging \textit{Basketball} and \textit{Balance Hard} tasks.

From Text to Forecasts: Bridging Modality Gap with Temporal Evolution Semantic Space

arXiv:2603.12664v1 Announce Type: cross Abstract: Incorporating textual information into time-series forecasting holds promise for addressing event-driven non-stationarity; however, a fundamental modality gap hinders effective fusion: textual descriptions express temporal impacts implicitly and qualitatively, whereas forecasting models rely on explicit and quantitative signals. Through controlled semi-synthetic experiments, we show that existing methods over-attend to redundant tokens and struggle to reliably translate textual semantics into usable numerical cues. To bridge this gap, we propose TESS, which introduces a Temporal Evolution Semantic Space as an intermediate bottleneck between modalities. This space consists of interpretable, numerically grounded temporal primitives (mean shift, volatility, shape, and lag) extracted from text by an LLM via structured prompting and filtered through confidence-aware gating. Experiments on four real-world datasets demonstrate up to a 29 percent reduction in forecasting error compared to state-of-the-art unimodal and multimodal baselines. The code will be released after acceptance.

NFATC2::NUTM2 Fusion Defines a Novel Primary Pulmonary Epithelial Tumor With a Distinctive Immunophenotype

Am J Surg Pathol. 2026 Jun 1;50(6):695-704. doi: 10.1097/PAS.0000000000002533. Epub 2026 Mar 13.

ABSTRACT

With the application of molecular techniques in pathologic diagnosis, several novel primary pulmonary epithelial tumors have been continuously discovered and classified under the WHO classification of thoracic tumors. Recently, a pulmonary tumor with NFATC2 :: NUTM2B fusion was first documented, but the spectrum of NFATC2::NUTM2 fusion variants and their associated pathologic features remains incompletely characterized. Coincidentally, we also found and described 6 primary pulmonary tumors harboring recurrent NFATC2::NUTM2A/E fusions through integrated genomic analysis. These patients, including 4 females and 2 males, with a median age of 53 years, presented with incidentally detected peripheral lung nodules composed of monotonous epithelioid cells arranged in cords, nests, and trabeculae within a prominent desmoplastic stroma. All tumors exhibited a consistent immunophenotype: CK5/6+/GATA3+/calponin+/EMA+/DOG1 (perinuclear dot-like staining)/p63-. High-throughput chromosome conformation capture (Hi-C) analysis showed the structural variation of NFATC2::NUTM2E in all 6 cases, whereas RNA sequencing detected the fusion transcripts in 5 cases ( NFATC2::NUTM2A , n=2; NFATC2::NUTM2E , n=3). Ultrastructural examination of 1 case suggested epithelial differentiation. All patients remained disease-free after complete resection (median follow-up: 24 mo; range: 9 to 41 mo). These findings define a novel primary pulmonary tumor entity driven by NFATC2::NUTM2 fusions, and characterized by a distinctive immunophenotype, expanding the spectrum of NUTM2 -associated neoplasms. Our study underscores the utility of multiomics approaches for characterizing rare neoplasms and provides a diagnostic framework for this entity.

PMID:41821426 | DOI:10.1097/PAS.0000000000002533

NFATC2::NUTM2 Fusion Defines a Novel Primary Pulmonary Epithelial Tumor With a Distinctive Immunophenotype

Am J Surg Pathol. 2026 Mar 13. doi: 10.1097/PAS.0000000000002533. Online ahead of print.

ABSTRACT

With the application of molecular techniques in pathologic diagnosis, several novel primary pulmonary epithelial tumors have been continuously discovered and classified under the WHO classification of thoracic tumors. Recently, a pulmonary tumor with NFATC2::NUTM2B fusion was first documented, but the spectrum of NFATC2::NUTM2 fusion variants and their associated pathologic features remains incompletely characterized. Coincidentally, we also found and described 6 primary pulmonary tumors harboring recurrent NFATC2::NUTM2A/E fusions through integrated genomic analysis. These patients, including 4 females and 2 males, with a median age of 53 years, presented with incidentally detected peripheral lung nodules composed of monotonous epithelioid cells arranged in cords, nests, and trabeculae within a prominent desmoplastic stroma. All tumors exhibited a consistent immunophenotype: CK5/6+/GATA3+/calponin+/EMA+/DOG1 (perinuclear dot-like staining)/p63-. High-throughput chromosome conformation capture (Hi-C) analysis showed the structural variation of NFATC2::NUTM2E in all 6 cases, whereas RNA sequencing detected the fusion transcripts in 5 cases (NFATC2::NUTM2A, n=2; NFATC2::NUTM2E, n=3). Ultrastructural examination of 1 case suggested epithelial differentiation. All patients remained disease-free after complete resection (median follow-up: 24 mo; range: 9 to 41 mo). These findings define a novel primary pulmonary tumor entity driven by NFATC2::NUTM2 fusions, and characterized by a distinctive immunophenotype, expanding the spectrum of NUTM2-associated neoplasms. Our study underscores the utility of multiomics approaches for characterizing rare neoplasms and provides a diagnostic framework for this entity.

PMID:41821426 | DOI:10.1097/PAS.0000000000002533

CSF1R T567M mutation induces microglial dysfunction and synaptic impairment in patient iPSC-derived cerebral organoids of CSF1R-related disorder

Cell Death Discovery, Published online: 12 March 2026; doi:10.1038/s41420-026-02995-2

CSF1R T567M mutation induces microglial dysfunction and synaptic impairment in patient iPSC-derived cerebral organoids of CSF1R-related disorder
❌