❌

Reading view

Let the Agent Steer: Closed-Loop Ranking Optimization via Influence Exchange

arXiv:2603.27765v2 Announce Type: replace Abstract: Recommendation ranking is fundamentally an influence allocation problem: a sorting formula distributes ranking influence among competing factors, and the business outcome depends on finding the optimal "exchange rates" among them. However, offline proxy metrics systematically misjudge how influence reallocation translates to online impact, with asymmetric bias across metrics that a single calibration factor cannot correct. We present Sortify, the first fully autonomous LLM-driven ranking optimization agent deployed in a large-scale production recommendation system. The agent reframes ranking optimization as continuous influence exchange, closing the full loop from diagnosis to parameter deployment without human intervention. It addresses structural problems through three mechanisms: (1) a dual-channel framework grounded in Savage's Subjective Expected Utility (SEU) that decouples offline-online transfer correction (Belief channel) from constraint penalty adjustment (Preference channel); (2) an LLM meta-controller operating on framework-level parameters rather than low-level search variables; (3) a persistent Memory DB with 7 relational tables for cross-round learning. Its core metric, Influence Share, provides a decomposable measure where all factor contributions sum to exactly 100%. Sortify has been deployed across two markets. In Country A, the agent pushed GMV from -3.6% to +9.2% within 7 rounds with peak orders reaching +12.5%. In Country B, a cold-start deployment achieved +4.15% GMV/UU and +3.58% Ads Revenue in a 7-day A/B test, leading to full production rollout.
  •  

InCoder-32B: Code Foundation Model for Industrial Scenarios

arXiv:2603.16790v3 Announce Type: replace-cross Abstract: Recent code large language models have achieved remarkable progress on general programming tasks. Nevertheless, their performance degrades significantly in industrial scenarios that require reasoning about hardware semantics, specialized language constructs, and strict resource constraints. To address these challenges, we introduce InCoder-32B (Industrial-Coder-32B), the first 32B-parameter code foundation model unifying code intelligence across chip design, GPU kernel optimization, embedded systems, compiler optimization, and 3D modeling. By adopting an efficient architecture, we train InCoder-32B from scratch with general code pre-training, curated industrial code annealing, mid-training that progressively extends context from 8K to 128K tokens with synthetic industrial reasoning data, and post-training with execution-grounded verification. We conduct extensive evaluation on 14 mainstream general code benchmarks and 9 industrial benchmarks spanning 4 specialized domains. Results show InCoder-32B achieves highly competitive performance on general tasks while establishing strong open-source baselines across industrial domains.
  •  

LDHA-driven lactate metabolism promotes MDSC activation and immunosuppressive microenvironment in prostate cancer

Oncogene, Published online: 01 April 2026; doi:10.1038/s41388-026-03737-5

LDHA-driven lactate metabolism promotes MDSC activation and immunosuppressive microenvironment in prostate cancer
  •  

A unified deep learning framework for cross-platform harmonization of multi-tracer PET quantification in neurodegenerative disease

npj Digital Medicine, Published online: 30 March 2026; doi:10.1038/s41746-026-02570-0

A unified deep learning framework for cross-platform harmonization of multi-tracer PET quantification in neurodegenerative disease
  •  

Targeting sialic acid metabolism: a therapeutic strategy against gastric cancer driven by WZ35

Cell Oncol (Dordr). 2026 Mar 23;49(2):60. doi: 10.1007/s13402-026-01194-6.

ABSTRACT

Glycolytic reprogramming is closely associated with the occurrence and progression of gastric cancer. Specifically, the energy derived from glucose metabolism and the cellular proteins by its intermediate products influence gastric cancer development. However, as an important branch of glucose metabolism, sialic acid metabolism and its mediated sialylation modifications remain insufficiently studied in gastric cancer, and their specific relationship with malignant tumor progression requires further exploration. This study employed a multi‑omics approach, integrating metabolomics, single‑cell RNA sequencing, and bulk RNA sequencing analyses, to investigate the metabolic landscape of gastric cancer and its associated alterations. The results indicated that sialic acid is a characteristic metabolite in malignant gastric cancer tissues. It modulates biological functions such as immune response, proliferative activity, and metabolic remodeling within gastric cancer tissues by influencing sialylation modifications. Furthermore, we identified the drug WZ35, which can inhibit the malignant proliferation of gastric cancer by targeting both sialic acid metabolism and sialylated protein modifications. We put forward a conjecture that the metabolism and modification of sialic acid promote the malignant development of gastric cancer, and we discovered that the drug WZ35 has an inhibitory effect on the sialic acid metabolism of gastric cancer.

GRAPHICAL ABSTRACT:

PMID:41870836 | PMC:PMC13009457 | DOI:10.1007/s13402-026-01194-6

  •  

MERIT: Memory-Enhanced Retrieval for Interpretable Knowledge Tracing

arXiv:2603.22289v1 Announce Type: cross Abstract: Knowledge Tracing (KT) models students' evolving knowledge states to predict future performance, serving as a foundation for personalized education. While traditional deep learning models achieve high accuracy, they often lack interpretability. Large Language Models (LLMs) offer strong reasoning capabilities but struggle with limited context windows and hallucinations. Furthermore, existing LLM-based methods typically require expensive fine-tuning, limiting scalability and adaptability to new data. We propose MERIT (Memory-Enhanced Retrieval for Interpretable Knowledge Tracing), a training-free framework combining frozen LLM reasoning with structured pedagogical memory. Rather than updating parameters, MERIT transforms raw interaction logs into an interpretable memory bank. The framework uses semantic denoising to categorize students into latent cognitive schemas and constructs a paradigm bank where representative error patterns are analyzed offline to generate explicit Chain-of-Thought (CoT) rationales. During inference, a hierarchical routing mechanism retrieves relevant contexts, while a logic-augmented module applies semantic constraints to calibrate predictions. By grounding the LLM in interpretable memory, MERIT achieves state-of-the-art performance on real-world datasets without gradient updates. This approach reduces computational costs and supports dynamic knowledge updates, improving the accessibility and transparency of educational diagnosis.
  •  

Mind Your HEARTBEAT! Claw Background Execution Inherently Enables Silent Memory Pollution

arXiv:2603.23064v2 Announce Type: cross Abstract: We identify a critical security vulnerability in mainstream Claw personal AI agents: untrusted content encountered during heartbeat-driven background execution can silently pollute agent memory and subsequently influence user-facing behavior without the user's awareness. This vulnerability arises from an architectural design shared across the Claw ecosystem: heartbeat background execution runs in the same session as user-facing conversation, so content ingested from any external source monitored in the background (including email, message channels, news feeds, code repositories, and social platforms) can enter the same memory context used for foreground interaction, often with limited user visibility and without clear source provenance. We formalize this process as an Exposure (E) $\rightarrow$ Memory (M) $\rightarrow$ Behavior (B) pathway: misinformation encountered during heartbeat execution enters the agent's short-term session context, potentially gets written into long-term memory, and later shapes downstream user-facing behavior. We instantiate this pathway in an agent-native social setting using MissClaw, a controlled research replica of Moltbook. We find that (1) social credibility cues, especially perceived consensus, are the dominant driver of short-term behavioral influence, with misleading rates up to 61%; (2) routine memory-saving behavior can promote short-term pollution into durable long-term memory at rates up to 91%, with cross-session behavioral influence reaching 76%; (3) under naturalistic browsing with content dilution and context pruning, pollution still crosses session boundaries. Overall, prompt injection is not required: ordinary social misinformation is sufficient to silently shape agent memory and behavior under heartbeat-driven background execution.
  •  

Dataset Distillation-based Hybrid Federated Learning on Non-IID Data

arXiv:2409.17517v3 Announce Type: replace-cross Abstract: In federated learning, the heterogeneity of client data has a great impact on the performance of model training. Many heterogeneity issues in this process are raised by non-independently and identically distributed (non-IID) data. To address the issue of label distribution skew, we propose a hybrid federated learning framework called HFLDD, which integrates dataset distillation to generate approximately independent and equally distributed (IID) data, thereby improving the performance of model training. In particular, we partition the clients into heterogeneous clusters, where the data labels among different clients within a cluster are unbalanced while the data labels among different clusters are balanced. The cluster heads collect distilled data from the corresponding cluster members, and conduct model training in collaboration with the server. This training process is like traditional federated learning on IID data, and hence effectively alleviates the impact of non-IID data on model training. We perform a comprehensive analysis of the convergence behavior, communication overhead, and computational complexity of the proposed HFLDD. Extensive experimental results based on multiple public datasets demonstrate that when data labels are severely imbalanced, the proposed HFLDD outperforms the baseline methods in terms of both test accuracy and communication cost.
  •  

FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control

arXiv:2603.12612v1 Announce Type: cross Abstract: Scaling Maximum Entropy Reinforcement Learning (RL) to high-dimensional humanoid control remains a formidable challenge, as the ``curse of dimensionality'' induces severe exploration inefficiency and training instability in expansive action spaces. Consequently, recent high-throughput paradigms have largely converged on deterministic policy gradients combined with massive parallel simulation. We challenge this compromise with FastDSAC, a framework that effectively unlocks the potential of maximum entropy stochastic policies for complex continuous control. We introduce Dimension-wise Entropy Modulation (DEM) to dynamically redistribute the exploration budget and enforce diversity, alongside a continuous distributional critic tailored to ensure value fidelity and mitigate high-dimensional value overestimation. Extensive evaluations on HumanoidBench and other continuous control tasks demonstrate that rigorously designed stochastic policies can consistently match or outperform deterministic baselines, achieving notable gains of 180\% and 400\% on the challenging \textit{Basketball} and \textit{Balance Hard} tasks.
  •  

From Text to Forecasts: Bridging Modality Gap with Temporal Evolution Semantic Space

arXiv:2603.12664v1 Announce Type: cross Abstract: Incorporating textual information into time-series forecasting holds promise for addressing event-driven non-stationarity; however, a fundamental modality gap hinders effective fusion: textual descriptions express temporal impacts implicitly and qualitatively, whereas forecasting models rely on explicit and quantitative signals. Through controlled semi-synthetic experiments, we show that existing methods over-attend to redundant tokens and struggle to reliably translate textual semantics into usable numerical cues. To bridge this gap, we propose TESS, which introduces a Temporal Evolution Semantic Space as an intermediate bottleneck between modalities. This space consists of interpretable, numerically grounded temporal primitives (mean shift, volatility, shape, and lag) extracted from text by an LLM via structured prompting and filtered through confidence-aware gating. Experiments on four real-world datasets demonstrate up to a 29 percent reduction in forecasting error compared to state-of-the-art unimodal and multimodal baselines. The code will be released after acceptance.
  •  

NFATC2::NUTM2 Fusion Defines a Novel Primary Pulmonary Epithelial Tumor With a Distinctive Immunophenotype

Am J Surg Pathol. 2026 Jun 1;50(6):695-704. doi: 10.1097/PAS.0000000000002533. Epub 2026 Mar 13.

ABSTRACT

With the application of molecular techniques in pathologic diagnosis, several novel primary pulmonary epithelial tumors have been continuously discovered and classified under the WHO classification of thoracic tumors. Recently, a pulmonary tumor with NFATC2 :: NUTM2B fusion was first documented, but the spectrum of NFATC2::NUTM2 fusion variants and their associated pathologic features remains incompletely characterized. Coincidentally, we also found and described 6 primary pulmonary tumors harboring recurrent NFATC2::NUTM2A/E fusions through integrated genomic analysis. These patients, including 4 females and 2 males, with a median age of 53 years, presented with incidentally detected peripheral lung nodules composed of monotonous epithelioid cells arranged in cords, nests, and trabeculae within a prominent desmoplastic stroma. All tumors exhibited a consistent immunophenotype: CK5/6+/GATA3+/calponin+/EMA+/DOG1 (perinuclear dot-like staining)/p63-. High-throughput chromosome conformation capture (Hi-C) analysis showed the structural variation of NFATC2::NUTM2E in all 6 cases, whereas RNA sequencing detected the fusion transcripts in 5 cases ( NFATC2::NUTM2A , n=2; NFATC2::NUTM2E , n=3). Ultrastructural examination of 1 case suggested epithelial differentiation. All patients remained disease-free after complete resection (median follow-up: 24 mo; range: 9 to 41 mo). These findings define a novel primary pulmonary tumor entity driven by NFATC2::NUTM2 fusions, and characterized by a distinctive immunophenotype, expanding the spectrum of NUTM2 -associated neoplasms. Our study underscores the utility of multiomics approaches for characterizing rare neoplasms and provides a diagnostic framework for this entity.

PMID:41821426 | DOI:10.1097/PAS.0000000000002533

  •  

NFATC2::NUTM2 Fusion Defines a Novel Primary Pulmonary Epithelial Tumor With a Distinctive Immunophenotype

Am J Surg Pathol. 2026 Mar 13. doi: 10.1097/PAS.0000000000002533. Online ahead of print.

ABSTRACT

With the application of molecular techniques in pathologic diagnosis, several novel primary pulmonary epithelial tumors have been continuously discovered and classified under the WHO classification of thoracic tumors. Recently, a pulmonary tumor with NFATC2::NUTM2B fusion was first documented, but the spectrum of NFATC2::NUTM2 fusion variants and their associated pathologic features remains incompletely characterized. Coincidentally, we also found and described 6 primary pulmonary tumors harboring recurrent NFATC2::NUTM2A/E fusions through integrated genomic analysis. These patients, including 4 females and 2 males, with a median age of 53 years, presented with incidentally detected peripheral lung nodules composed of monotonous epithelioid cells arranged in cords, nests, and trabeculae within a prominent desmoplastic stroma. All tumors exhibited a consistent immunophenotype: CK5/6+/GATA3+/calponin+/EMA+/DOG1 (perinuclear dot-like staining)/p63-. High-throughput chromosome conformation capture (Hi-C) analysis showed the structural variation of NFATC2::NUTM2E in all 6 cases, whereas RNA sequencing detected the fusion transcripts in 5 cases (NFATC2::NUTM2A, n=2; NFATC2::NUTM2E, n=3). Ultrastructural examination of 1 case suggested epithelial differentiation. All patients remained disease-free after complete resection (median follow-up: 24 mo; range: 9 to 41 mo). These findings define a novel primary pulmonary tumor entity driven by NFATC2::NUTM2 fusions, and characterized by a distinctive immunophenotype, expanding the spectrum of NUTM2-associated neoplasms. Our study underscores the utility of multiomics approaches for characterizing rare neoplasms and provides a diagnostic framework for this entity.

PMID:41821426 | DOI:10.1097/PAS.0000000000002533

  •  

CSF1R T567M mutation induces microglial dysfunction and synaptic impairment in patient iPSC-derived cerebral organoids of CSF1R-related disorder

Cell Death Discovery, Published online: 12 March 2026; doi:10.1038/s41420-026-02995-2

CSF1R T567M mutation induces microglial dysfunction and synaptic impairment in patient iPSC-derived cerebral organoids of CSF1R-related disorder
  •  

Regulatory mechanisms of ALKBH5/CIITA axis in the synergistic modulation of hepatocellular carcinoma radiotherapy and immunotherapy

Genes Immun. 2026 Mar 10. doi: 10.1038/s41435-026-00382-6. Online ahead of print.

ABSTRACT

The prognosis for hepatocellular carcinoma remains grim. Combining radiotherapy with immune checkpoint blockade (ICB) has shown potential to enhance therapeutic outcomes, yet there is a pressing need for further advancements. Our previous research demonstrated that this combined approach suppresses ALKBH5 gene expression and increases m6A modification levels in hepatocellular carcinoma tissues. High-throughput sequencing and detailed molecular analysis revealed that inhibiting ALKBH5 amplifies CIITA m6A modifications post-therapy. This modulation triggers MHC II molecule expression in tumors, facilitating the presentation of tumor-associated antigens to CD4 + T lymphocytes and the recruitment of CD8 + T cells for an anti-tumor immune response. Building on these findings, we engineered a CIITA vector with a specific site mutation to confirm that the regulation of CIITA by the combined radiotherapy and immunotherapy is mediated through m6A methylation. Consequently, we established a comprehensive network involving ALKBH5, CIITA, MHC II, and CD4+ and CD8 + T cells. To elucidate the role and underlying molecular mechanisms of this combined therapy in reshaping the tumor immune microenvironment for hepatocellular carcinoma, we employed multi-omics approaches across in vitro, animal model, and clinical multi-dimensional studies, offering novel insights for enhancing treatment efficacy.

PMID:41807814 | DOI:10.1038/s41435-026-00382-6

  •  

SynPlanResearch-R1: Encouraging Tool Exploration for Deep Research with Synthetic Plans

arXiv:2603.07853v1 Announce Type: new Abstract: Research Agents enable models to gather information from the web using tools to answer user queries, requiring them to dynamically interleave internal reasoning with tool use. While such capabilities can in principle be learned via reinforcement learning with verifiable rewards (RLVR), we observe that agents often exhibit poor exploration behaviors, including premature termination and biased tool usage. As a result, RLVR alone yields limited improvements. We propose SynPlanResearch-R1, a framework that synthesizes tool-use trajectories that encourage deeper exploration to shape exploration during cold-start supervised fine-tuning, providing a strong initialization for subsequent RL. Across seven multi-hop and open-web benchmarks, \framework improves performance by up to 6.0% on Qwen3-8B and 5.8% on Qwen3-4B backbones respectively compared to SOTA baselines. Further analyses of tool-use patterns and training dynamics compared to baselines shed light on the factors underlying these gains. Our code is publicly available at https://github.com/HansiZeng/syn-plan-research.
  •  

Contextual Counterfactual Credit Assignment for Multi-Agent Reinforcement Learning in LLM Collaboration

arXiv:2603.06859v1 Announce Type: cross Abstract: Cooperative multi-agent reinforcement learning (MARL) systems powered by large language models (LLMs) are frequently optimized via sparse terminal-only feedback. This shared signal entangles upstream decisions, obstructing accurate decision-level credit assignment. To address this trajectory-level diffusion, we introduce Contextual Counterfactual Credit Assignment (\textbf{\texttt{C3}}). Instead of distributing rewards across an entire episode, \textbf{\texttt{C3}} isolates the causal impact of individual messages by freezing the exact transcript-derived context, evaluating context-matched alternatives via fixed-continuation replay, and applying a leave-one-out (LOO) baseline. This localized intervention extracts unbiased, low-variance marginal advantages for standard policy-gradient optimization. Evaluated across five mathematical and coding benchmarks under matched budgets, \textbf{\texttt{C3}} improves terminal performance over established baselines. Mechanistic diagnostics further show that these gains are accompanied by higher credit fidelity, lower contextual variance, and stronger inter-agent causal dependence. Our code is available at https://github.com/EIT-EAST-Lab/C3.
  •  

MetaWorld-X: Hierarchical World Modeling via VLM-Orchestrated Experts for Humanoid Loco-Manipulation

arXiv:2603.08572v1 Announce Type: cross Abstract: Learning natural, stable, and compositionally generalizable whole-body control policies for humanoid robots performing simultaneous locomotion and manipulation (loco-manipulation) remains a fundamental challenge in robotics. Existing reinforcement learning approaches typically rely on a single monolithic policy to acquire multiple skills, which often leads to cross-skill gradient interference and motion pattern conflicts in high-degree-of-freedom systems. As a result, generated behaviors frequently exhibit unnatural movements, limited stability, and poor generalization to complex task compositions. To address these limitations, we propose MetaWorld-X, a hierarchical world model framework for humanoid control. Guided by a divide-and-conquer principle, our method decomposes complex control problems into a set of specialized expert policies (Specialized Expert Policies, SEP). Each expert is trained under human motion priors through imitation-constrained reinforcement learning, introducing biomechanically consistent inductive biases that ensure natural and physically plausible motion generation. Building upon this foundation, we further develop an Intelligent Routing Mechanism (IRM) supervised by a Vision-Language Model (VLM), enabling semantic-driven expert composition. The VLM-guided router dynamically integrates expert policies according to high-level task semantics, facilitating compositional generalization and adaptive execution in multi-stage loco-manipulation tasks.
  •  

M4Diffuser: Multi-View Diffusion Policy with Manipulability-Aware Control for Robust Mobile Manipulation

arXiv:2509.14980v2 Announce Type: replace-cross Abstract: Mobile manipulation requires the coordinated control of a mobile base and a robotic arm while simultaneously perceiving both global scene context and fine-grained object details. Existing single-view approaches often fail in unstructured environments due to limited fields of view, exploration, and generalization abilities. Moreover, classical controllers, although stable, struggle with efficiency and manipulability near singularities. To address these challenges, we propose M4Diffuser, a hybrid framework that integrates a Multi-View Diffusion Policy with a novel Reduced and Manipulability-aware QP (ReM-QP) controller for mobile manipulation. The diffusion policy leverages proprioceptive states and complementary camera perspectives with both close-range object details and global scene context to generate task-relevant end-effector goals in the world frame. These high-level goals are then executed by the ReM-QP controller, which eliminates slack variables for computational efficiency and incorporates manipulability-aware preferences for robustness near singularities. Comprehensive experiments in simulation and real-world environments show that M4Diffuser achieves 7 to 56 percent higher success rates and reduces collisions by 3 to 31 percent over baselines. Our approach demonstrates robust performance for smooth whole-body coordination, and strong generalization to unseen tasks, paving the way for reliable mobile manipulation in unstructured environments. Details of the demo and supplemental material are available on our project website https://sites.google.com/view/m4diffuser.
  •  

TRIM27-controlled endothelium-derived exosomes play a central role in podocyte injury in diabetic kidney disease

Cell Death Discovery, Published online: 07 March 2026; doi:10.1038/s41420-026-02953-y

TRIM27-controlled endothelium-derived exosomes play a central role in podocyte injury in diabetic kidney disease
  •  
❌