❌

Normal view

Pan-cancer oncolytic virotherapy through disruption of tumor cell mitochondrial dynamics

Li and colleagues identified RhoA as a “redox rheostat” governing mitochondrial dynamics during oncolytic virotherapy and thereby engineered rNDV-RHOA, an NDV-based oncolytic virus overexpressing RhoA. This tumor-targeted RhoA overexpression synergizes oxidative stress and viral oncolysis, transcending conventional oncolysis by surmounting tumor heterogeneity through exploiting inherent tumor redox dependency.

JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition

arXiv:2609.10451v1 Announce Type: new Abstract: Real-world GUI usage frequently involves workflows that span multiple devices and platforms, requiring the transfer of intermediate results, maintenance of shared state, and coordination across heterogeneous environments. However, existing GUI benchmarks overwhelmingly evaluate agents on single-device, statically defined tasks, thus leaving such cross-device capabilities largely unexamined, resulting in an overly optimistic assessment of agents' readiness for real-world usage. We introduce JarvisGUI, a dynamic benchmark that evaluates GUI agents on cross-device workflows requiring coordinated interaction across heterogeneous platforms, including Android, Windows, and Ubuntu. Specifically, JarvisGUI formulates GUI tasks as input-output transformations under a lightweight type system, which allows us to automatically compose multi-step, cross-device workflows and dynamically evaluate agent performance within a unified framework. By evaluating agents in virtual environments spanning multiple operating systems, JarvisGUI reveals that state-of-the-art open-source GUI agents struggle with the state-transfer awareness, cross-platform contextual reasoning, and long-horizon dependency management required for real-world workflows, exposing a critical capability gap invisible to existing benchmarks.

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

arXiv:2609.04298v2 Announce Type: replace Abstract: Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of 8 models spanning capability tiers across 54 benchmarks; every model is run with Terminus-2 and with one of 3 native harnesses. This enables a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduce Harbor-Index, a curated set of 82 difficult, diverse, and high-quality tasks spanning 29 benchmarks, refined from the adapted suite through difficulty filtering, AI and human audit, and an audit-and-fix loop. Harbor-Index preserves the challenge and breadth of large-scale agentic evaluations while being affordable to run; no evaluated model-harness configuration exceeds 30% pass rate, and the strongest (GPT-5.5 with Codex) reaches 28.0%. We release the adapters, evaluation results, in-depth analysis, and Harbor-Index as open-source artifacts to support more reliable and comprehensive evaluation of language-model agents.

RAU: Reference-based Anatomical Understanding with Vision Language Models

arXiv:2509.22404v2 Announce Type: replace-cross Abstract: Anatomical understanding, which is the ability to identify, localize, or segment anatomical structures, is critical in medical image analysis; however, its progress is constrained by the scarcity of expert-labeled data. A promising remedy is to leverage an annotated reference image to guide the interpretation of an unlabeled target. Although recent vision-language models (VLMs) exhibit non-trivial visual reasoning, their reference-based understanding and fine-grained localization remain limited. We introduce RAU, a framework for reference-based anatomical understanding with VLMs. We first show that a VLM learns to identify anatomical regions through relative spatial reasoning between reference and target images, trained on a moderately sized dataset. We validate this capability through visual question answering (VQA) and bounding box prediction. Next, we demonstrate that the VLM-derived spatial cues can be seamlessly integrated with the fine-grained segmentation capability of SAM2, enabling localization and pixel-level segmentation of small anatomical regions, such as vessel segments. Across two in-distribution and two out-of-distribution datasets, RAU consistently outperforms a SAM2 fine-tuning baseline using the same memory setup, yielding more accurate segmentations and more reliable localization. More importantly, its generalization ability to unseen modalities makes it scalable to unseen datasets, a property crucial for medical image applications. To the best of our knowledge, RAU is the first to explore the capability of VLMs for reference-based identification, localization, and segmentation of anatomical structures in medical images. Its promising performance highlights the potential of VLM-driven approaches for anatomical understanding in automated clinical workflows.

S1P-TREM2 axis protects immunosuppressive neutrophils from ferroptosis to promote tumour progression in hepatocellular carcinoma

Gut. 2026 Sep 7:gutjnl-2025-337414. doi: 10.1136/gutjnl-2025-337414. Online ahead of print.

ABSTRACT

BACKGROUND: Neutrophils are increasingly recognised as immunosuppressive drivers of hepatocellular carcinoma (HCC), yet their persistence in the oxidative, lipid-rich tumour microenvironment remains poorly understood.

OBJECTIVE: To elucidate the metabolic and molecular programmes that enable tumour-associated neutrophils (TANs) to resist ferroptosis and sustain immunosuppression in HCC.

DESIGN: We employed human HCC samples, multiple murine HCC models, transcriptomic and lipidomic profiling, genetic loss-of-function systems and therapeutic interventions. Ferroptosis sensitivity, lipid metabolic rewiring and immunological consequences of TANs were systematically evaluated across models and validated in patient datasets and biospecimens.

RESULTS: TANs in human HCC and mouse models exhibit pronounced lipid accumulation and oxidative stress compared with peripheral neutrophils. Multi-omic profiling revealed that TANs are enriched for lipid-binding gene programmes and undergo rewiring towards sphingolipid and unsaturated fatty acid metabolism. We identified triggering receptor expressed on myeloid cells 2 (TREM2) as a key lipid-sensing receptor selectively expressed in TANs. Functional deletion of TREM2 reprogrammed the tumour immune microenvironment, restoring CD8+ T cell activity and suppressing HCC progression. Mechanistically, tumour-derived sphingosine-1-phosphate (S1P) activates TREM2, triggering nuclear factor erythroid 2-related factor 2 (NRF2)-mediated transcription of glutathione peroxidase 4 (GPX4) and solute carrier family 7 member 11 (SLC7A11), thereby promoting ferroptosis resistance. TREM2 expression is transcriptionally induced by granulocyte-macrophage colony-stimulating factor-signal transducer and activator of transcription 3 (GM-CSF-STAT3) signalling. Genetic deletion of TREM2, clustered regularly interspaced short palindromic repeats/CRISPR-associated protein 9 (CRISPR/Cas9)-mediated knockout of sphingosine kinase 1/2 (SPHK1/2) in tumour cells, or pharmacological inhibition of S1P synthesis disrupts this protective lipid-immune circuit, sensitises TANs to ferroptosis and restricts tumour growth. Therapeutically, a peptide-based TREM2 inhibitor reprogrammes TANs, restores CD8+ T cell function and enhances anti-programmed cell death protein 1 (PD-1) immunotherapy efficacy. Clinically, TREM2+ polymorphonuclear myeloid-derived suppressor cells (PMN-MDSCs) are enriched in HCC tumours, correlate with SPHK1/2 expression and T cell dysfunction and associate with poor patient prognosis.

CONCLUSION: Our study uncovers the S1P-TREM2-NRF2 axis as a critical metabolic-immune circuit that preserves neutrophil survival and immunosuppressive function in HCC. Targeting this lipid-dependent ferroptosis resistance pathway offers a promising therapeutic strategy to overcome immunotherapy resistance in liver cancer.

PMID:42705697 | DOI:10.1136/gutjnl-2025-337414

S1P-TREM2 axis protects immunosuppressive neutrophils from ferroptosis to promote tumour progression in hepatocellular carcinoma

Gut. 2026 Sep 7:gutjnl-2025-337414. doi: 10.1136/gutjnl-2025-337414. Online ahead of print.

ABSTRACT

BACKGROUND: Neutrophils are increasingly recognised as immunosuppressive drivers of hepatocellular carcinoma (HCC), yet their persistence in the oxidative, lipid-rich tumour microenvironment remains poorly understood.

OBJECTIVE: To elucidate the metabolic and molecular programmes that enable tumour-associated neutrophils (TANs) to resist ferroptosis and sustain immunosuppression in HCC.

DESIGN: We employed human HCC samples, multiple murine HCC models, transcriptomic and lipidomic profiling, genetic loss-of-function systems and therapeutic interventions. Ferroptosis sensitivity, lipid metabolic rewiring and immunological consequences of TANs were systematically evaluated across models and validated in patient datasets and biospecimens.

RESULTS: TANs in human HCC and mouse models exhibit pronounced lipid accumulation and oxidative stress compared with peripheral neutrophils. Multi-omic profiling revealed that TANs are enriched for lipid-binding gene programmes and undergo rewiring towards sphingolipid and unsaturated fatty acid metabolism. We identified triggering receptor expressed on myeloid cells 2 (TREM2) as a key lipid-sensing receptor selectively expressed in TANs. Functional deletion of TREM2 reprogrammed the tumour immune microenvironment, restoring CD8+ T cell activity and suppressing HCC progression. Mechanistically, tumour-derived sphingosine-1-phosphate (S1P) activates TREM2, triggering nuclear factor erythroid 2-related factor 2 (NRF2)-mediated transcription of glutathione peroxidase 4 (GPX4) and solute carrier family 7 member 11 (SLC7A11), thereby promoting ferroptosis resistance. TREM2 expression is transcriptionally induced by granulocyte-macrophage colony-stimulating factor-signal transducer and activator of transcription 3 (GM-CSF-STAT3) signalling. Genetic deletion of TREM2, clustered regularly interspaced short palindromic repeats/CRISPR-associated protein 9 (CRISPR/Cas9)-mediated knockout of sphingosine kinase 1/2 (SPHK1/2) in tumour cells, or pharmacological inhibition of S1P synthesis disrupts this protective lipid-immune circuit, sensitises TANs to ferroptosis and restricts tumour growth. Therapeutically, a peptide-based TREM2 inhibitor reprogrammes TANs, restores CD8+ T cell function and enhances anti-programmed cell death protein 1 (PD-1) immunotherapy efficacy. Clinically, TREM2+ polymorphonuclear myeloid-derived suppressor cells (PMN-MDSCs) are enriched in HCC tumours, correlate with SPHK1/2 expression and T cell dysfunction and associate with poor patient prognosis.

CONCLUSION: Our study uncovers the S1P-TREM2-NRF2 axis as a critical metabolic-immune circuit that preserves neutrophil survival and immunosuppressive function in HCC. Targeting this lipid-dependent ferroptosis resistance pathway offers a promising therapeutic strategy to overcome immunotherapy resistance in liver cancer.

PMID:42705697 | DOI:10.1136/gutjnl-2025-337414

Integrative Multi-Omics Analysis Identifies Thrombosis-Associated Molecular Features Linked to Germline Susceptibility and Immune Cell Communication in Gastric Cancer

Chem Biol Drug Des. 2026 Sep;108(3):e70386. doi: 10.1111/cbdd.70386.

ABSTRACT

Emerging evidence indicates that coagulation-related molecular programs are associated with thrombosis, tumor progression, and molecular dysregulation in gastric cancer (GC). However, thrombosis-associated molecular features in GC and their potential links to inherited susceptibility remain insufficiently understood. Integrated analyses of transcriptomic data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) datasets were performed to identify thrombosis-associated genes and establish a machine learning-based prognostic signature. Genome-wide association study (GWAS), expression quantitative trait loci (eQTL), transcriptome-wide association study (TWAS), and Mendelian randomization (MR) analyses were conducted to investigate susceptibility-associated transcriptional programs in GC. Functional assays were used to evaluate candidate genes associated with malignant phenotypes. Single-cell RNA sequencing (scRNA-seq) and cell-cell communication analyses were further performed to characterize cell-type-specific expression patterns and potential intercellular interactions. A total of 22 differentially expressed thrombosis-associated genes were identified, and a prognostic signature comprising 14 genes was established. The signature stratified patients into high- and low-risk groups and showed prognostic performance in both the training and validation cohorts. Integrative GWAS, eQTL, and TWAS analyses identified susceptibility-associated transcriptional programs that were positively correlated with the thrombosis-associated risk score. Silencing ACTN2 and CRYAB significantly reduced GC cell migration and invasion. scRNA-seq analysis revealed relatively high CRYAB expression in neutrophils, and CellChat analysis suggested potential neutrophil-B cell interactions involving COLLAGEN-related signaling. This integrative multi-omics study identified a thrombosis-associated molecular signature linked to prognosis and germline susceptibility-associated transcriptional programs in GC. ACTN2 and CRYAB may represent candidate genes associated with GC cell migration and invasion, while single-cell analysis suggested potential immune-related communication features.

PMID:42681916 | PMC:PMC13534880 | DOI:10.1111/cbdd.70386

DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning

arXiv:2605.23939v1 Announce Type: new Abstract: Web agents require both high-level reasoning (for task decomposition) and low-level interactions (for page elements manipulation) to conduct different tasks. However, these knowledge types differ fundamentally: reasoning knowledge (e.g., booking a flight requires first searching for routes) is abstract and transferable across websites, while interaction knowledge (e.g., clicking the Search button at a specific coordinate on Site A) depends heavily on page-specific contexts. Existing methods store experiences uniformly. This creates a dilemma: abstract representations lose executability on concrete pages, while concrete representations fail to generalize across domains. This entanglement limits capability accumulation: on new websites, agents either fail to recognize reusable task logic due to surface-level differences or attempt infeasible actions from outdated page structures. To disentangle them, we propose DRIVE, a dual-level skill modeling framework separating historical experience into natural language reasoning skills, which capture transferable task logic, and programmatic interaction skills, grounding abstract actions to executable operations. A scene-aware coordination mechanism adaptively retrieves and invokes these dual-level skills based on task semantics. DRIVE also uses skill-level reflection to identify hierarchy-specific failure modes, enabling targeted skill library expansion and refinement. Experiments across five WebArena domains show DRIVE attains an average task success rate of 52.8%, exceeding the skill-free baseline by 7.3 percentage points. Further ablations show reasoning and interaction skills provide distinct, complementary benefits, supporting separation of transferable task logic from executable page-level operations.

SPACE: Unifying Symmetric and Asymmetric Routing Problems for Generalist Neural Solver

arXiv:2605.24484v1 Announce Type: new Abstract: Generalist neural routing solvers have shown great potential in solving diverse vehicle routing problems (VRPs) with a unified model. However, existing solvers are typically limited to symmetric settings or degrade in performance when switching to asymmetric settings due to input inconsistencies or inherent structural differences, substantially limiting their practicality in real-world scenarios that encompass both scenarios. To address this limitation, we define the spatial position of each node based on the relative distances to a specific set of pivots and further propose a Spatial Pivot-Aligned Coordinate-free Embedding (SPACE) framework that unifies node representation and solution generation across symmetric and asymmetric VRPs. Specifically, we construct a bidirectional Frechet representation using a novel furthest pivot sampling strategy to enable invariant node representations across distinct problem settings. Furthermore, we introduce a weight-decomposed adaptive decoding mechanism that decouples geometric perception from problem representations, mitigating the overfitting of constraint decisions to a specific geometry setting. Extensive experiments on 110 VRP variants, comprising 55 symmetric problems and their asymmetric counterparts, demonstrate that SPACE achieves promising zero-shot generalization in both symmetric and asymmetric VRPs.

NeurIPS: Neuro-anatomical Inductive Priors for Sphere-based Brain Decoding

arXiv:2605.24993v1 Announce Type: new Abstract: Current fMRI decoders face a performance-fidelity trade-off where efficient ID encoders outperform geometrically faithful surface-based models. We argue this is partly driven by inefficient surface tokenization and the failure to use anatomy as a predictive signal. We present NeurIPS, a framework that improves surface-based decoding by reframing anatomical variation from a nuisance to a powerful inductive prior. NeurIPS unites two innovations: a Selective ROI Spherical Tokenizer (SRST) for efficient geometric encoding, and a Structure-Guided Mixture of Experts (SG-MoE) that explicitly models individual anatomy using cortical features. On the Natural Scenes Dataset, NeurIPS establishes a new state-of-the-art for surface decoders and achieves performance comparable to strong 1D baselines. This is achieved with unprecedented efficiency, as the model converges dramatically faster (10 vs. 600 epochs). This efficiency enables rapid adaptation to new subjects using only 20% of data and ensures robust scalability as the training cohort is expanded. Ablations provide causal evidence that these gains are driven by the model's use of cortical features, not by memorizing subject IDs. By leveraging anatomical priors, NeurIPS provides a principled and scalable path toward robust, generalizable brain decoding.

Agent-Centric Social Trajectory Prediction: A Free Energy Principle Perspective

arXiv:2605.25748v1 Announce Type: new Abstract: Trajectory prediction methods have demonstrated remarkable capabilities in capturing complex motion patterns. However, existing methods rely on global state assumptions, suffer from insufficient belief inference under partial observability, and lack cognitive behavioral constraints in prediction. These limitations severely compromise both deployment feasibility and physical plausibility in real-world settings. In this work, we propose FEP-Diff, an agent-centric trajectory prediction framework grounded in the Free Energy Principle, aimed at achieving cognitively plausible predictions under realistic constraints. Specifically, a dual-branch spatiotemporal encoder extracts ego-motion dynamics and social interaction cues from local observations. Building upon this, a goal-conditioned belief learner infers multimodal latent belief distributions optimized via a free-energy objective, with a social consistency constraint on the local neighborhood graph to promote cognitive alignment among neighboring agents. Finally, a residual diffusion trajectory generator is conditioned on the learned belief representations with token-level proxy conditioning, producing precise and diverse future predictions. Extensive experiments on five public benchmarks demonstrate that FEP-Diff consistently outperforms state-of-the-art methods under restricted observability. Code: https://anonymous.4open.science/r/FEP-Diff-8876.

CITYREP: A Unified Benchmark for Urban Representations Across Cities, Tasks, and Modalities

arXiv:2605.26036v1 Announce Type: new Abstract: Urban representation learning encodes complex urban environments into general-purpose embeddings for diverse downstream tasks and emerging urban foundation models. However, current evaluations are limited, typically focusing on one or two cities and tasks and relying on random splits that introduce spatial leakage, leading to inflated performance and weak support for cross-location generalization and fair comparison. To address this, we propose CityRep, a unified benchmark that evaluates urban representations across data modalities, cities, and tasks using spatially structured splits. CityRep consists of three key components: (1) a spatial unit-agnostic evaluation framework that supports heterogeneous urban representations through a standardized alignment module; (2) a unified evaluation protocol using block-based spatial splits to mitigate spatial leakage and enable rigorous model comparison; and (3) an extensible multi-city, multi-task benchmark suite spanning 8 cities and 8 tasks across regression, classification, and distribution prediction. We evaluate 11 representative urban representation models. Results show that performance is highly sensitive to the split protocol, with random splits inflating scores and altering model rankings. We also observe substantial variability across cities and tasks, underscoring the need for generalization-aware evaluation. CityRep is released as a reproducible benchmark with datasets, evaluation pipelines, and diagnostic tools to facilitate fair comparison and support future research in urban representation learning towards urban foundation models.

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models

arXiv:2605.24011v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute make deployment on edge platforms impractical. Aggressive, sub-4-bit weight quantization is the natural solution, yet existing post-training quantization (PTQ) methods suffer severe performance degradation in this regime. To address this, we introduce ActQuant, an action-guided mixed-precision PTQ framework that operates in two stages: (1) an inter-tensor bit allocator that assigns each weight matrix a single bit-width based on how much it contributes to predicting the agent's actions; (2) an intra-tensor scale optimizer tunes per-block quantization scales using action-aware curvature, so that dynamic range is concentrated on the weights most influential for control. To deliver the on-device benefits of our aggressive quantization, we further introduce OmniModel.cpp, an agentic conversion pipeline that ports architectures into a native C/C++ runtime with efficient low-bit kernels. We evaluate ActQuant both in simulation and on a real-world 6-DoF UR3 arm, with all models deployed through OmniModel.cpp. On the LIBERO benchmark, ActQuant is the only method that operates at or below 3 bits-per-weight, retaining 95.0% on OpenVLA-OFT and 94.8% on $\pi_{0.5}$. Pushed further, ActQuant reaches 2.5 bpw at 90.1% on OpenVLA-OFT, compressing the backbone from 14.3 GB to 2.7 GB (5.3$\times$). On the physical UR3 arm, $\pi_{0.5}$ quantized with ActQuant retains the baseline's success rate while reducing the memory footprint by 2.5$\times$.

Rethinking Federated Unlearning via the Lens of Memorization

arXiv:2605.24545v1 Announce Type: cross Abstract: Federated learning (FL) increasingly needs machine unlearning to comply with privacy regulations. However, existing federated unlearning approaches may overlook the overlapping information between the unlearning and remaining data, leading to ineffective unlearning and unfairness between clients. In this work, we revisit federated unlearning through the lens of memorization. We argue that unlearning should mainly remove the unique memorized information attributable to the data to be forgotten, while preserving overlapping patterns that are also supported by the remaining data. Specifically, we propose Grouped Memorization Evaluation, an example-level metric that separates memorized knowledge from overlapping knowledge. Building on this metric, we introduce Federated Memorization Pruning (FedMemPrune), a pruning-based unlearning approach that resets redundant parameters responsible for memorization. Extensive experiments show that FedMemPrune closely matches retraining-based unlearning baselines while more effectively eliminating memorization than existing federated unlearning algorithms, yielding strong unlearning performance without sacrificing the utility of retained knowledge.

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation

arXiv:2605.24675v1 Announce Type: cross Abstract: Translating text embedded in Web images is crucial for improving content accessibility and cross-lingual information retrieval, particularly within social media and e-commerce domains. Although Large Vision-Language Models (LVLMs) have advanced multimodal understanding, applying them to Web image translation remains challenging due to the visual representation gap: standard encoders often prioritize high-level semantics over the fine-grained visual details required for recognizing diverse character morphologies. To address this challenge, we propose VaaWIT, an end-to-end framework that adapts Large Language Models for multilingual Web image translation. The framework introduces two key technical contributions: (1) a Dual-Stream Attention Module (DSAM), which facilitates bidirectional interaction between multilingual semantic features and detailed visual representations, thereby synthesizing unified features robust to textual variations; and (2) a Visual-Aware Adapter (VAA), a parameter-efficient fine-tuning strategy that dynamically injects these fused visual cues into the frozen LLM backbone. This design enables the model to align the visual context with linguistic reasoning effectively while minimizing computational costs. Extensive experiments on eight tasks on three public benchmarks demonstrate that VaaWIT significantly outperforms state-of-the-art (SOTA) open-source baselines and achieves competitive performance against proprietary models. These results validate the efficacy of integrating fine-grained visual perception into LLMs for complex Web content analysis.

RealBench: Benchmarking Data-Driven Numerical Weather Forecasting Under Operational Conditions and Extreme Event Challenges

arXiv:2605.24945v1 Announce Type: cross Abstract: Accurate evaluation of weather forecasting models is critical for their reliable deployment in real-world applications. However, existing benchmarks predominantly rely on reanalysis products such as ERA5, which are generated through delayed data assimilation and do not reflect the constraints of real-time operational forecasting, thereby resulting in a systematic mismatch between benchmark performance and real-world forecasting. In this work, we introduce RealBench, a next-generation benchmark for AI weather forecasting that emphasizes realistic evaluation under operational conditions. RealBench features a strictly out-of-distribution test set spanning 2025 to eliminate data leakage and capture recent atmospheric regimes. It integrates multiple data sources, including low-latency operational analysis and a large-scale global in-situ observation dataset comprising over 10,000 stations, enabling direct evaluation against real atmospheric measurements. Beyond standard global metrics, RealBench provides a comprehensive evaluation framework for high-impact extreme events, including heatwaves, cold surges, and tropical cyclones, using event-specific metrics that better reflect real-world forecasting priorities. The evaluation results reveal substantial discrepancies between reanalysis-based metrics and real-world performance, particularly concerning extreme events. By highlighting the limitations of existing benchmarks, this work establishes a more faithful and operationally relevant evaluation paradigm, providing a rigorous foundation for advancing next-generation AI weather forecasting systems. The benchmark implementation is available at: https://github.com/lixruize-del/NWP-Benchmark.

DeGRe: Dense-supervised Generative Reranking for Recommendation

arXiv:2605.25749v1 Announce Type: cross Abstract: In multi-stage recommender systems, reranking optimizes overall utility by capturing intra-list contextual dependencies, yet its central challenge lies in exploring optimal sequences within an exponentially large permutation space. Recent studies have shifted towards end-to-end generative frameworks, which typically leverage list-wise rewards or preference alignment to guide generator training. However, these methods still face two critical issues. First is the heuristic label bias. Existing methods often construct training targets based on simple rules, such as promoting clicked items to the top, while ignoring causal dependencies within the list context. Second is the credit assignment problem. Sparse list-level posterior rewards fail to directly guide intermediate steps in sequence generation, leading to ambiguous optimization directions. To address these issues, we propose DeGRe (Dense-supervised Generative Reranking), a generative reranking framework that bridges the gap between offline exploration and online efficiency through dense supervision. The core of DeGRe lies in its offline-online decoupled design. During the offline phase, we introduce a Lookahead Evaluator based on cumulative regression, which leverages beam search to actively mine high-value lookahead sequences in the unexposed space. During training, we transform the step-wise value estimations from the evaluator into dense supervision signals and distill them into a lightweight Online Generator. This mechanism enables the generator to internalize lookahead planning capabilities, requiring only a single efficient greedy decoding pass during online inference to approximate the global optimum. Experiments demonstrate that DeGRe outperforms baseline models on public benchmarks and industrial datasets. We have successfully deployed DeGRe on Taobao Flash Shopping, significantly improving online recommendations.

Agent Learning via Early Experience

arXiv:2510.08558v3 Announce Type: replace Abstract: A long-term goal of language agents is to learn and improve through their own experience, ultimately outperforming humans in complex, real-world tasks. However, training agents from experience data with reinforcement learning remains difficult in many environments, which either lack verifiable rewards (e.g., websites) or require inefficient long-horizon rollouts (e.g., multi-turn tool use). As a result, most current agents rely on supervised fine-tuning on expert data, which is challenging to scale and generalizes poorly. This limitation stems from the nature of expert demonstrations: they capture only a narrow range of scenarios, and expose the agent to limited environment diversity. We address this limitation with a middle-ground paradigm we call early experience: interaction data generated by the agent's own actions, where the resulting future states serve as supervision without reward signals. Within this paradigm, we study two strategies of using such data: (1) implicit world modeling, which uses collected states to ground the policy in environment dynamics; and (2) self-reflection, where the agent learns from its suboptimal actions to improve reasoning and decision-making. Evaluation across eight diverse environments and multiple model families shows that our approaches consistently improve effectiveness and out-of-domain generalization, highlighting the value of early experience. Moreover, in environments with verifiable rewards, our results provide promising signals that early experience offers a strong foundation for subsequent reinforcement learning, making it a practical bridge between imitation learning and fully experience-driven agents.

PathMem: Toward Cognition-Aligned Memory Transformation for Pathology MLLMs

arXiv:2603.09943v2 Announce Type: replace Abstract: Computational pathology demands both visual pattern recognition and dynamic integration of structured domain knowledge, including taxonomy, grading criteria, and clinical evidence. In practice, diagnostic reasoning requires linking morphological evidence with formal diagnostic and grading criteria. Although multimodal large language models (MLLMs) demonstrate strong vision language reasoning capabilities, they lack explicit mechanisms for structured knowledge integration and interpretable memory control. As a result, existing models struggle to consistently incorporate pathology-specific diagnostic standards during reasoning. Inspired by the hierarchical memory process of human pathologists, we propose PathMem, a memory-centric multimodal framework for pathology MLLMs. PathMem organizes structured pathology knowledge as a long-term memory (LTM) and introduces a Memory Transformer that models the dynamic transition from LTM to working memory (WM) through multimodal memory activation and context-aware knowledge grounding, enabling context-aware memory refinement for downstream reasoning. PathMem achieves SOTA performance across benchmarks, improving WSI-Bench report generation (12.8% WSI-Precision, 10.1% WSI-Relevance) and open-ended diagnosis by 9.7% and 8.9% over prior WSI-based models.

UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents

arXiv:2604.11557v2 Announce Type: replace Abstract: Tool-use capability is a fundamental component of LLM agents, enabling them to interact with external systems through structured function calls. However, existing research exhibits inconsistent interaction representations, largely overlooks the structural distribution of tool-use trajectories, and relies on incompatible evaluation benchmarks. We present UniToolCall, a unified framework for tool learning that standardizes the entire pipeline from toolset construction and dataset generation to evaluation. The framework curates a large tool pool of 22k+ tools and constructs a hybrid training corpus of 390k+ instances by combining 10 standardized public datasets with structurally controlled synthetic trajectories. It explicitly models diverse interaction patterns, including single-hop vs. multi-hop and single-turn vs. multi-turn, while capturing both serial and parallel execution structures. To support coherent multi-turn reasoning, we further introduce an Anchor Linkage mechanism that enforces cross-turn dependencies. Furthermore, we convert 7 public benchmarks into a unified Query--Action--Observation--Answer (QAOA) representation with fine-grained evaluation at the function-call, turn, and conversation levels. Experiments show that fine-tuning Qwen3-8B on our dataset substantially improves tool-use performance. Under the distractor-heavy Hybrid-20 setting, achieves 93.0% single-turn Strict Precision, outperforming commercial models including GPT, Gemini, and Claude.
❌