❌

Normal view

GLARE: Generative Learning via Adversarial Reward Estimation For Social Dynamics Forecasting

arXiv:2609.12165v1 Announce Type: new Abstract: Meeting continuation requires tracking the agenda, speaker roles, participant intentions, and disagreement across long multi-party discussions. We introduce the Meeting Dynamic Forecasting Benchmark (MDFB), constructed from 2,207 real-world meetings and 24,794 future-facing queries. Given a transcript prefix and an active question, a model generates a plausible multi-turn continuation in one call. We evaluate utility---progress toward the question---and human-likeness---plausible conversational flow and role consistency---without requiring exact reproduction of the observed future. We further present GLARE, an adaptation of adversarial imitation learning to conditional language generation. A discriminator ranks the observed continuation above samples from the current actor, and its score supplies a KL-regularized policy reward; retraining on current-policy negatives allows the reward landscape to evolve with the actor. GLARE attains average human-evaluated win rates of 0.66 on utility and 0.70 on human-likeness, outperforming SFT and SPIN while remaining below the observed human continuation. We also demonstrate MDFB as a social reasoning arena for comparing general-purpose models, including closed-source systems, through reference-assisted judgments. Together, these studies illustrate the benchmark's use for both task-specific learning and output-based evaluation of meeting behavior.

LifeFuse-Mem: Lifecycle-Aware State Fusion Against Temporary Overwriting for Long-Term Memory

arXiv:2609.12436v1 Announce Type: new Abstract: Long-running LLM agents require memory mechanisms that maintain coherent internal states across interactions. We study a lifecycle-labeled memory setting in which write episodes provide lifecycle metadata during training, and phase-aware readout is used during evaluation. This setting reflects the need to distinguish information that should remain influential across future interactions from information that should affect only the current context. A mismatch between these lifecycles can cause temporary information to overwrite durable knowledge, leading to behavioral drift in persistent agents. Within this setting, we introduce \textbf{LifeFuse-Mem}, a lifecycle-aware neural memory framework that separates information according to its temporal commitment. LifeFuse-Mem uses dedicated memory components and lifecycle-aware updates to allow stable and transient knowledge to evolve locally without converting temporary context into durable state. On the controlled anti-overwrite benchmark, LifeFuse-Mem improves acquisition-controlled retention and reduces temporary overwrite; on two public long-memory benchmarks, it remains broadly competitive. These results suggest that explicit lifecycle signals can help diagnose and mitigate overwrite in compact online memory.

Beyond Generation and Accuracy: Diagnosing and Enhancing Visual Chain-of-Thought for Geometry Problem Solving

arXiv:2609.12606v1 Announce Type: new Abstract: While multimodal reasoning has advanced rapidly, solving complex geometry problems critically hinges on active visual assistance, such as constructing auxiliary lines, spurring the rise of Visual Chain-of-Thought (VCoT). However, existing evaluations typically assess visual generation quality and final answer accuracy in isolation, failing to examine whether intermediate visual aids are geometrically valid, effectively utilized in subsequent reasoning, or causally responsible for task success. To bridge this gap, we introduce GeoVAD-Bench, a diagnostic benchmark that pairs a fine-grained five-dimensional trajectory diagnosis covering perception, auxiliary quality, utilization, deductive reasoning, and final correctness with controlled No-Aux, Auto-Aux, and GT-Aux intervention settings to systematically isolate intermediate error modes, the causal gains of visual aids, and the resulting autonomy gap. Our findings reveal that while high-quality auxiliary aids offer substantial theoretical gains for geometric problem solving, autonomous generation is frequently hampered by compounding errors across geometric perception, faithful visual manipulation, visual-state grounding, and deductive reasoning. Guided by these diagnostic insights, we establish a specialized data construction pipeline encompassing geometric perception, diagram editing, and interleaved visual-textual reasoning trajectories, and develop a progressive SFT and multimodal RL training framework. The resulting model, GeoWeave-8B, outperforms the base model by +25.3% in final geometric accuracy and achieves a +30.4% gain in process average across the four intermediate diagnostic dimensions.

MedCollab: IBIS-Guided Multi-Agent Collaboration with Hierarchical Disease Relation Chains for Clinical Diagnosis

arXiv:2603.01131v4 Announce Type: replace-cross Abstract: Clinical diagnosis is a gradual process of evidence integration, in which physicians move from symptoms and medical history to examinations, competing hypotheses, disease relations, and treatment decisions. Large language models have advanced medical text understanding and generation. Yet their clinical use remains limited by weak evidence grounding, opaque reasoning, and inconsistent links among differential diagnosis, final diagnosis, diagnostic basis, and treatment planning. We introduce MedCollab, a multi-agent framework for full-cycle clinical diagnosis and report generation. MedCollab coordinates specialist and examination agents according to patient records. It structures agent deliberation with an Issue-Based Information System (IBIS) protocol, so that each diagnostic position is supported by patient-specific evidence and medical knowledge. It also builds Hierarchical Disease Relation Chains (HDRC) to connect accepted hypotheses through progression, complication, and comorbidity relations. During multi-round deliberation, a verifier-guided consensus module evaluates evidence support, medical plausibility, and logical conflicts. It then adjusts agent contributions and filters unsupported reasoning. Experiments on ClinicalBench and MIMIC-IV show that MedCollab outperforms leading LLMs and medical multi-agent baselines in diagnostic accuracy, evidence consistency, and clinical reasoning quality. These results indicate that structured and auditable collaboration can produce more faithful and clinically coherent diagnostic reports.

Narrative review of the staging classification controversy in stage N3 small cell lung cancer: from the perspective of overlapping Veterans Administration Lung Study Group and International Association for the Study of Lung Cancer definitions

J Thorac Dis. 2026 Aug 31;18(8):950. doi: 10.21037/jtd-2026-1704. Epub 2026 Aug 28.

ABSTRACT

BACKGROUND AND OBJECTIVE: Traditionally, two primary systems have been employed for staging small cell lung cancer (SCLC): the Veterans Administration Lung Study Group (VALG) system and the International Association for the Study of Lung Cancer (IASLC) tumor, node, metastasis (TNM) system. The term "limited disease" is defined differently: VALG characterizes it as disease encompassed within a single tolerable radiation field, while IASLC defines it as the lack of distant metastases (M0). Patients with N3 disease frequently satisfy VALG extensive-stage (ES) criteria while meeting IASLC limited-stage (LS) criteria, resulting in a notable staging discrepancy. Therefore, this review aims to clarify the clinical challenges posed by this staging overlap and provide insights for standardizing staging terminology and optimizing therapeutic decision-making in N3 SCLC.

METHODS: A narrative review utilizing a systematized search strategy was conducted. While strict adherence to PRISMA guidelines was not pursued because the extensive heterogeneity of the literature precluded a formal meta-analysis, rigorous search criteria were applied to minimize selection bias. Databases including PubMed, Web of Science, Embase, the Cochrane Library, and China National Knowledge Infrastructure (CNKI) were searched for literature from January 2000 to March 2026. Studies examining stage N3 SCLC, spatial metastatic burden, and definitional inconsistencies between the VALG and IASLC staging systems were analyzed to assess their effects on treatment dosimetry, systemic therapy, and survival outcomes.

KEY CONTENT AND FINDINGS: The staging overlap in N3 SCLC leads to heterogeneous clinical management depending on its spatial metastatic burden, and this highly variable cohort can be stratified into distinct prognostic subgroups based on the anatomical distribution (single-region vs. multi-region) of the involved lymph nodes.

CONCLUSIONS: These findings should guide clinical trial design and terminology. Clinical decision-making must transcend historical paradigms and technical constraints. Future strategies must incorporate spatial evaluations of metastatic burden alongside innovative multimodal tools, such as artificial intelligence (AI) and multi-omics, to facilitate tailored therapy for SCLC.

PMID:42724560 | PMC:PMC13559235 | DOI:10.21037/jtd-2026-1704

Narrative review of the staging classification controversy in stage N3 small cell lung cancer: from the perspective of overlapping Veterans Administration Lung Study Group and International Association for the Study of Lung Cancer definitions

J Thorac Dis. 2026 Aug 31;18(8):950. doi: 10.21037/jtd-2026-1704. Epub 2026 Aug 28.

ABSTRACT

BACKGROUND AND OBJECTIVE: Traditionally, two primary systems have been employed for staging small cell lung cancer (SCLC): the Veterans Administration Lung Study Group (VALG) system and the International Association for the Study of Lung Cancer (IASLC) tumor, node, metastasis (TNM) system. The term "limited disease" is defined differently: VALG characterizes it as disease encompassed within a single tolerable radiation field, while IASLC defines it as the lack of distant metastases (M0). Patients with N3 disease frequently satisfy VALG extensive-stage (ES) criteria while meeting IASLC limited-stage (LS) criteria, resulting in a notable staging discrepancy. Therefore, this review aims to clarify the clinical challenges posed by this staging overlap and provide insights for standardizing staging terminology and optimizing therapeutic decision-making in N3 SCLC.

METHODS: A narrative review utilizing a systematized search strategy was conducted. While strict adherence to PRISMA guidelines was not pursued because the extensive heterogeneity of the literature precluded a formal meta-analysis, rigorous search criteria were applied to minimize selection bias. Databases including PubMed, Web of Science, Embase, the Cochrane Library, and China National Knowledge Infrastructure (CNKI) were searched for literature from January 2000 to March 2026. Studies examining stage N3 SCLC, spatial metastatic burden, and definitional inconsistencies between the VALG and IASLC staging systems were analyzed to assess their effects on treatment dosimetry, systemic therapy, and survival outcomes.

KEY CONTENT AND FINDINGS: The staging overlap in N3 SCLC leads to heterogeneous clinical management depending on its spatial metastatic burden, and this highly variable cohort can be stratified into distinct prognostic subgroups based on the anatomical distribution (single-region vs. multi-region) of the involved lymph nodes.

CONCLUSIONS: These findings should guide clinical trial design and terminology. Clinical decision-making must transcend historical paradigms and technical constraints. Future strategies must incorporate spatial evaluations of metastatic burden alongside innovative multimodal tools, such as artificial intelligence (AI) and multi-omics, to facilitate tailored therapy for SCLC.

PMID:42724560 | PMC:PMC13559235 | DOI:10.21037/jtd-2026-1704

A clinically-oriented foundation model for intraoperative pathology

Nature Medicine, Published online: 10 September 2026; doi:10.1038/s41591-026-04703-0

CRISP, a vision-based pathology foundation model developed exclusively from frozen section slides, supports treatment decision-making throughout the surgical workflow with superior performance to current foundation models and extensive validation, including in a prospective cohort.

Pan-cancer oncolytic virotherapy through disruption of tumor cell mitochondrial dynamics

Li and colleagues identified RhoA as a “redox rheostat” governing mitochondrial dynamics during oncolytic virotherapy and thereby engineered rNDV-RHOA, an NDV-based oncolytic virus overexpressing RhoA. This tumor-targeted RhoA overexpression synergizes oxidative stress and viral oncolysis, transcending conventional oncolysis by surmounting tumor heterogeneity through exploiting inherent tumor redox dependency.

JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition

arXiv:2609.10451v1 Announce Type: new Abstract: Real-world GUI usage frequently involves workflows that span multiple devices and platforms, requiring the transfer of intermediate results, maintenance of shared state, and coordination across heterogeneous environments. However, existing GUI benchmarks overwhelmingly evaluate agents on single-device, statically defined tasks, thus leaving such cross-device capabilities largely unexamined, resulting in an overly optimistic assessment of agents' readiness for real-world usage. We introduce JarvisGUI, a dynamic benchmark that evaluates GUI agents on cross-device workflows requiring coordinated interaction across heterogeneous platforms, including Android, Windows, and Ubuntu. Specifically, JarvisGUI formulates GUI tasks as input-output transformations under a lightweight type system, which allows us to automatically compose multi-step, cross-device workflows and dynamically evaluate agent performance within a unified framework. By evaluating agents in virtual environments spanning multiple operating systems, JarvisGUI reveals that state-of-the-art open-source GUI agents struggle with the state-transfer awareness, cross-platform contextual reasoning, and long-horizon dependency management required for real-world workflows, exposing a critical capability gap invisible to existing benchmarks.

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

arXiv:2609.04298v2 Announce Type: replace Abstract: Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of 8 models spanning capability tiers across 54 benchmarks; every model is run with Terminus-2 and with one of 3 native harnesses. This enables a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduce Harbor-Index, a curated set of 82 difficult, diverse, and high-quality tasks spanning 29 benchmarks, refined from the adapted suite through difficulty filtering, AI and human audit, and an audit-and-fix loop. Harbor-Index preserves the challenge and breadth of large-scale agentic evaluations while being affordable to run; no evaluated model-harness configuration exceeds 30% pass rate, and the strongest (GPT-5.5 with Codex) reaches 28.0%. We release the adapters, evaluation results, in-depth analysis, and Harbor-Index as open-source artifacts to support more reliable and comprehensive evaluation of language-model agents.

RAU: Reference-based Anatomical Understanding with Vision Language Models

arXiv:2509.22404v2 Announce Type: replace-cross Abstract: Anatomical understanding, which is the ability to identify, localize, or segment anatomical structures, is critical in medical image analysis; however, its progress is constrained by the scarcity of expert-labeled data. A promising remedy is to leverage an annotated reference image to guide the interpretation of an unlabeled target. Although recent vision-language models (VLMs) exhibit non-trivial visual reasoning, their reference-based understanding and fine-grained localization remain limited. We introduce RAU, a framework for reference-based anatomical understanding with VLMs. We first show that a VLM learns to identify anatomical regions through relative spatial reasoning between reference and target images, trained on a moderately sized dataset. We validate this capability through visual question answering (VQA) and bounding box prediction. Next, we demonstrate that the VLM-derived spatial cues can be seamlessly integrated with the fine-grained segmentation capability of SAM2, enabling localization and pixel-level segmentation of small anatomical regions, such as vessel segments. Across two in-distribution and two out-of-distribution datasets, RAU consistently outperforms a SAM2 fine-tuning baseline using the same memory setup, yielding more accurate segmentations and more reliable localization. More importantly, its generalization ability to unseen modalities makes it scalable to unseen datasets, a property crucial for medical image applications. To the best of our knowledge, RAU is the first to explore the capability of VLMs for reference-based identification, localization, and segmentation of anatomical structures in medical images. Its promising performance highlights the potential of VLM-driven approaches for anatomical understanding in automated clinical workflows.

S1P-TREM2 axis protects immunosuppressive neutrophils from ferroptosis to promote tumour progression in hepatocellular carcinoma

Gut. 2026 Sep 7:gutjnl-2025-337414. doi: 10.1136/gutjnl-2025-337414. Online ahead of print.

ABSTRACT

BACKGROUND: Neutrophils are increasingly recognised as immunosuppressive drivers of hepatocellular carcinoma (HCC), yet their persistence in the oxidative, lipid-rich tumour microenvironment remains poorly understood.

OBJECTIVE: To elucidate the metabolic and molecular programmes that enable tumour-associated neutrophils (TANs) to resist ferroptosis and sustain immunosuppression in HCC.

DESIGN: We employed human HCC samples, multiple murine HCC models, transcriptomic and lipidomic profiling, genetic loss-of-function systems and therapeutic interventions. Ferroptosis sensitivity, lipid metabolic rewiring and immunological consequences of TANs were systematically evaluated across models and validated in patient datasets and biospecimens.

RESULTS: TANs in human HCC and mouse models exhibit pronounced lipid accumulation and oxidative stress compared with peripheral neutrophils. Multi-omic profiling revealed that TANs are enriched for lipid-binding gene programmes and undergo rewiring towards sphingolipid and unsaturated fatty acid metabolism. We identified triggering receptor expressed on myeloid cells 2 (TREM2) as a key lipid-sensing receptor selectively expressed in TANs. Functional deletion of TREM2 reprogrammed the tumour immune microenvironment, restoring CD8+ T cell activity and suppressing HCC progression. Mechanistically, tumour-derived sphingosine-1-phosphate (S1P) activates TREM2, triggering nuclear factor erythroid 2-related factor 2 (NRF2)-mediated transcription of glutathione peroxidase 4 (GPX4) and solute carrier family 7 member 11 (SLC7A11), thereby promoting ferroptosis resistance. TREM2 expression is transcriptionally induced by granulocyte-macrophage colony-stimulating factor-signal transducer and activator of transcription 3 (GM-CSF-STAT3) signalling. Genetic deletion of TREM2, clustered regularly interspaced short palindromic repeats/CRISPR-associated protein 9 (CRISPR/Cas9)-mediated knockout of sphingosine kinase 1/2 (SPHK1/2) in tumour cells, or pharmacological inhibition of S1P synthesis disrupts this protective lipid-immune circuit, sensitises TANs to ferroptosis and restricts tumour growth. Therapeutically, a peptide-based TREM2 inhibitor reprogrammes TANs, restores CD8+ T cell function and enhances anti-programmed cell death protein 1 (PD-1) immunotherapy efficacy. Clinically, TREM2+ polymorphonuclear myeloid-derived suppressor cells (PMN-MDSCs) are enriched in HCC tumours, correlate with SPHK1/2 expression and T cell dysfunction and associate with poor patient prognosis.

CONCLUSION: Our study uncovers the S1P-TREM2-NRF2 axis as a critical metabolic-immune circuit that preserves neutrophil survival and immunosuppressive function in HCC. Targeting this lipid-dependent ferroptosis resistance pathway offers a promising therapeutic strategy to overcome immunotherapy resistance in liver cancer.

PMID:42705697 | DOI:10.1136/gutjnl-2025-337414

S1P-TREM2 axis protects immunosuppressive neutrophils from ferroptosis to promote tumour progression in hepatocellular carcinoma

Gut. 2026 Sep 7:gutjnl-2025-337414. doi: 10.1136/gutjnl-2025-337414. Online ahead of print.

ABSTRACT

BACKGROUND: Neutrophils are increasingly recognised as immunosuppressive drivers of hepatocellular carcinoma (HCC), yet their persistence in the oxidative, lipid-rich tumour microenvironment remains poorly understood.

OBJECTIVE: To elucidate the metabolic and molecular programmes that enable tumour-associated neutrophils (TANs) to resist ferroptosis and sustain immunosuppression in HCC.

DESIGN: We employed human HCC samples, multiple murine HCC models, transcriptomic and lipidomic profiling, genetic loss-of-function systems and therapeutic interventions. Ferroptosis sensitivity, lipid metabolic rewiring and immunological consequences of TANs were systematically evaluated across models and validated in patient datasets and biospecimens.

RESULTS: TANs in human HCC and mouse models exhibit pronounced lipid accumulation and oxidative stress compared with peripheral neutrophils. Multi-omic profiling revealed that TANs are enriched for lipid-binding gene programmes and undergo rewiring towards sphingolipid and unsaturated fatty acid metabolism. We identified triggering receptor expressed on myeloid cells 2 (TREM2) as a key lipid-sensing receptor selectively expressed in TANs. Functional deletion of TREM2 reprogrammed the tumour immune microenvironment, restoring CD8+ T cell activity and suppressing HCC progression. Mechanistically, tumour-derived sphingosine-1-phosphate (S1P) activates TREM2, triggering nuclear factor erythroid 2-related factor 2 (NRF2)-mediated transcription of glutathione peroxidase 4 (GPX4) and solute carrier family 7 member 11 (SLC7A11), thereby promoting ferroptosis resistance. TREM2 expression is transcriptionally induced by granulocyte-macrophage colony-stimulating factor-signal transducer and activator of transcription 3 (GM-CSF-STAT3) signalling. Genetic deletion of TREM2, clustered regularly interspaced short palindromic repeats/CRISPR-associated protein 9 (CRISPR/Cas9)-mediated knockout of sphingosine kinase 1/2 (SPHK1/2) in tumour cells, or pharmacological inhibition of S1P synthesis disrupts this protective lipid-immune circuit, sensitises TANs to ferroptosis and restricts tumour growth. Therapeutically, a peptide-based TREM2 inhibitor reprogrammes TANs, restores CD8+ T cell function and enhances anti-programmed cell death protein 1 (PD-1) immunotherapy efficacy. Clinically, TREM2+ polymorphonuclear myeloid-derived suppressor cells (PMN-MDSCs) are enriched in HCC tumours, correlate with SPHK1/2 expression and T cell dysfunction and associate with poor patient prognosis.

CONCLUSION: Our study uncovers the S1P-TREM2-NRF2 axis as a critical metabolic-immune circuit that preserves neutrophil survival and immunosuppressive function in HCC. Targeting this lipid-dependent ferroptosis resistance pathway offers a promising therapeutic strategy to overcome immunotherapy resistance in liver cancer.

PMID:42705697 | DOI:10.1136/gutjnl-2025-337414

Integrative Multi-Omics Analysis Identifies Thrombosis-Associated Molecular Features Linked to Germline Susceptibility and Immune Cell Communication in Gastric Cancer

Chem Biol Drug Des. 2026 Sep;108(3):e70386. doi: 10.1111/cbdd.70386.

ABSTRACT

Emerging evidence indicates that coagulation-related molecular programs are associated with thrombosis, tumor progression, and molecular dysregulation in gastric cancer (GC). However, thrombosis-associated molecular features in GC and their potential links to inherited susceptibility remain insufficiently understood. Integrated analyses of transcriptomic data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) datasets were performed to identify thrombosis-associated genes and establish a machine learning-based prognostic signature. Genome-wide association study (GWAS), expression quantitative trait loci (eQTL), transcriptome-wide association study (TWAS), and Mendelian randomization (MR) analyses were conducted to investigate susceptibility-associated transcriptional programs in GC. Functional assays were used to evaluate candidate genes associated with malignant phenotypes. Single-cell RNA sequencing (scRNA-seq) and cell-cell communication analyses were further performed to characterize cell-type-specific expression patterns and potential intercellular interactions. A total of 22 differentially expressed thrombosis-associated genes were identified, and a prognostic signature comprising 14 genes was established. The signature stratified patients into high- and low-risk groups and showed prognostic performance in both the training and validation cohorts. Integrative GWAS, eQTL, and TWAS analyses identified susceptibility-associated transcriptional programs that were positively correlated with the thrombosis-associated risk score. Silencing ACTN2 and CRYAB significantly reduced GC cell migration and invasion. scRNA-seq analysis revealed relatively high CRYAB expression in neutrophils, and CellChat analysis suggested potential neutrophil-B cell interactions involving COLLAGEN-related signaling. This integrative multi-omics study identified a thrombosis-associated molecular signature linked to prognosis and germline susceptibility-associated transcriptional programs in GC. ACTN2 and CRYAB may represent candidate genes associated with GC cell migration and invasion, while single-cell analysis suggested potential immune-related communication features.

PMID:42681916 | PMC:PMC13534880 | DOI:10.1111/cbdd.70386

DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning

arXiv:2605.23939v1 Announce Type: new Abstract: Web agents require both high-level reasoning (for task decomposition) and low-level interactions (for page elements manipulation) to conduct different tasks. However, these knowledge types differ fundamentally: reasoning knowledge (e.g., booking a flight requires first searching for routes) is abstract and transferable across websites, while interaction knowledge (e.g., clicking the Search button at a specific coordinate on Site A) depends heavily on page-specific contexts. Existing methods store experiences uniformly. This creates a dilemma: abstract representations lose executability on concrete pages, while concrete representations fail to generalize across domains. This entanglement limits capability accumulation: on new websites, agents either fail to recognize reusable task logic due to surface-level differences or attempt infeasible actions from outdated page structures. To disentangle them, we propose DRIVE, a dual-level skill modeling framework separating historical experience into natural language reasoning skills, which capture transferable task logic, and programmatic interaction skills, grounding abstract actions to executable operations. A scene-aware coordination mechanism adaptively retrieves and invokes these dual-level skills based on task semantics. DRIVE also uses skill-level reflection to identify hierarchy-specific failure modes, enabling targeted skill library expansion and refinement. Experiments across five WebArena domains show DRIVE attains an average task success rate of 52.8%, exceeding the skill-free baseline by 7.3 percentage points. Further ablations show reasoning and interaction skills provide distinct, complementary benefits, supporting separation of transferable task logic from executable page-level operations.

SPACE: Unifying Symmetric and Asymmetric Routing Problems for Generalist Neural Solver

arXiv:2605.24484v1 Announce Type: new Abstract: Generalist neural routing solvers have shown great potential in solving diverse vehicle routing problems (VRPs) with a unified model. However, existing solvers are typically limited to symmetric settings or degrade in performance when switching to asymmetric settings due to input inconsistencies or inherent structural differences, substantially limiting their practicality in real-world scenarios that encompass both scenarios. To address this limitation, we define the spatial position of each node based on the relative distances to a specific set of pivots and further propose a Spatial Pivot-Aligned Coordinate-free Embedding (SPACE) framework that unifies node representation and solution generation across symmetric and asymmetric VRPs. Specifically, we construct a bidirectional Frechet representation using a novel furthest pivot sampling strategy to enable invariant node representations across distinct problem settings. Furthermore, we introduce a weight-decomposed adaptive decoding mechanism that decouples geometric perception from problem representations, mitigating the overfitting of constraint decisions to a specific geometry setting. Extensive experiments on 110 VRP variants, comprising 55 symmetric problems and their asymmetric counterparts, demonstrate that SPACE achieves promising zero-shot generalization in both symmetric and asymmetric VRPs.

NeurIPS: Neuro-anatomical Inductive Priors for Sphere-based Brain Decoding

arXiv:2605.24993v1 Announce Type: new Abstract: Current fMRI decoders face a performance-fidelity trade-off where efficient ID encoders outperform geometrically faithful surface-based models. We argue this is partly driven by inefficient surface tokenization and the failure to use anatomy as a predictive signal. We present NeurIPS, a framework that improves surface-based decoding by reframing anatomical variation from a nuisance to a powerful inductive prior. NeurIPS unites two innovations: a Selective ROI Spherical Tokenizer (SRST) for efficient geometric encoding, and a Structure-Guided Mixture of Experts (SG-MoE) that explicitly models individual anatomy using cortical features. On the Natural Scenes Dataset, NeurIPS establishes a new state-of-the-art for surface decoders and achieves performance comparable to strong 1D baselines. This is achieved with unprecedented efficiency, as the model converges dramatically faster (10 vs. 600 epochs). This efficiency enables rapid adaptation to new subjects using only 20% of data and ensures robust scalability as the training cohort is expanded. Ablations provide causal evidence that these gains are driven by the model's use of cortical features, not by memorizing subject IDs. By leveraging anatomical priors, NeurIPS provides a principled and scalable path toward robust, generalizable brain decoding.

Agent-Centric Social Trajectory Prediction: A Free Energy Principle Perspective

arXiv:2605.25748v1 Announce Type: new Abstract: Trajectory prediction methods have demonstrated remarkable capabilities in capturing complex motion patterns. However, existing methods rely on global state assumptions, suffer from insufficient belief inference under partial observability, and lack cognitive behavioral constraints in prediction. These limitations severely compromise both deployment feasibility and physical plausibility in real-world settings. In this work, we propose FEP-Diff, an agent-centric trajectory prediction framework grounded in the Free Energy Principle, aimed at achieving cognitively plausible predictions under realistic constraints. Specifically, a dual-branch spatiotemporal encoder extracts ego-motion dynamics and social interaction cues from local observations. Building upon this, a goal-conditioned belief learner infers multimodal latent belief distributions optimized via a free-energy objective, with a social consistency constraint on the local neighborhood graph to promote cognitive alignment among neighboring agents. Finally, a residual diffusion trajectory generator is conditioned on the learned belief representations with token-level proxy conditioning, producing precise and diverse future predictions. Extensive experiments on five public benchmarks demonstrate that FEP-Diff consistently outperforms state-of-the-art methods under restricted observability. Code: https://anonymous.4open.science/r/FEP-Diff-8876.

CITYREP: A Unified Benchmark for Urban Representations Across Cities, Tasks, and Modalities

arXiv:2605.26036v1 Announce Type: new Abstract: Urban representation learning encodes complex urban environments into general-purpose embeddings for diverse downstream tasks and emerging urban foundation models. However, current evaluations are limited, typically focusing on one or two cities and tasks and relying on random splits that introduce spatial leakage, leading to inflated performance and weak support for cross-location generalization and fair comparison. To address this, we propose CityRep, a unified benchmark that evaluates urban representations across data modalities, cities, and tasks using spatially structured splits. CityRep consists of three key components: (1) a spatial unit-agnostic evaluation framework that supports heterogeneous urban representations through a standardized alignment module; (2) a unified evaluation protocol using block-based spatial splits to mitigate spatial leakage and enable rigorous model comparison; and (3) an extensible multi-city, multi-task benchmark suite spanning 8 cities and 8 tasks across regression, classification, and distribution prediction. We evaluate 11 representative urban representation models. Results show that performance is highly sensitive to the split protocol, with random splits inflating scores and altering model rankings. We also observe substantial variability across cities and tasks, underscoring the need for generalization-aware evaluation. CityRep is released as a reproducible benchmark with datasets, evaluation pipelines, and diagnostic tools to facilitate fair comparison and support future research in urban representation learning towards urban foundation models.
❌