❌

Normal view

RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases

arXiv:2609.10092v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as research agents, yet their ability to track shifts in research attention is difficult to evaluate because reviews and research ideas lack uniquely verifiable outcomes. We introduce Research Attention Prediction (RAP), a rolling benchmark covering 278 AI/ML fields and 1,390 episodes. At each cut-off, an LLM agent searches a temporally restricted arXiv corpus and predicts the next six months' paper shares across eight frozen research directions. Search generally helps, but all four diagnostic models perform worse than an exact-count exponentially weighted moving average (EWMA) baseline in compositional accuracy. We identify two linked bottlenecks. Under cumulative-history access, State carry-forward outperforms direct Forecast for all four diagnostic models; frozen-evidence replay links a shared component of this reversal to Forecast-oriented policies retrieving a smaller share of recent evidence. Even with exact historical activity, future-specific updating remains limited, with only GPT-5.5 plus reopened Search slightly surpassing EWMA. Fine-tuning on realised outcomes improves Qwen3-4B's forecast Spearman correlation by 0.105 on held-out fields at later origins, with gains also on change-rich episodes.

S1P-TREM2 axis protects immunosuppressive neutrophils from ferroptosis to promote tumour progression in hepatocellular carcinoma

Gut. 2026 Sep 7:gutjnl-2025-337414. doi: 10.1136/gutjnl-2025-337414. Online ahead of print.

ABSTRACT

BACKGROUND: Neutrophils are increasingly recognised as immunosuppressive drivers of hepatocellular carcinoma (HCC), yet their persistence in the oxidative, lipid-rich tumour microenvironment remains poorly understood.

OBJECTIVE: To elucidate the metabolic and molecular programmes that enable tumour-associated neutrophils (TANs) to resist ferroptosis and sustain immunosuppression in HCC.

DESIGN: We employed human HCC samples, multiple murine HCC models, transcriptomic and lipidomic profiling, genetic loss-of-function systems and therapeutic interventions. Ferroptosis sensitivity, lipid metabolic rewiring and immunological consequences of TANs were systematically evaluated across models and validated in patient datasets and biospecimens.

RESULTS: TANs in human HCC and mouse models exhibit pronounced lipid accumulation and oxidative stress compared with peripheral neutrophils. Multi-omic profiling revealed that TANs are enriched for lipid-binding gene programmes and undergo rewiring towards sphingolipid and unsaturated fatty acid metabolism. We identified triggering receptor expressed on myeloid cells 2 (TREM2) as a key lipid-sensing receptor selectively expressed in TANs. Functional deletion of TREM2 reprogrammed the tumour immune microenvironment, restoring CD8+ T cell activity and suppressing HCC progression. Mechanistically, tumour-derived sphingosine-1-phosphate (S1P) activates TREM2, triggering nuclear factor erythroid 2-related factor 2 (NRF2)-mediated transcription of glutathione peroxidase 4 (GPX4) and solute carrier family 7 member 11 (SLC7A11), thereby promoting ferroptosis resistance. TREM2 expression is transcriptionally induced by granulocyte-macrophage colony-stimulating factor-signal transducer and activator of transcription 3 (GM-CSF-STAT3) signalling. Genetic deletion of TREM2, clustered regularly interspaced short palindromic repeats/CRISPR-associated protein 9 (CRISPR/Cas9)-mediated knockout of sphingosine kinase 1/2 (SPHK1/2) in tumour cells, or pharmacological inhibition of S1P synthesis disrupts this protective lipid-immune circuit, sensitises TANs to ferroptosis and restricts tumour growth. Therapeutically, a peptide-based TREM2 inhibitor reprogrammes TANs, restores CD8+ T cell function and enhances anti-programmed cell death protein 1 (PD-1) immunotherapy efficacy. Clinically, TREM2+ polymorphonuclear myeloid-derived suppressor cells (PMN-MDSCs) are enriched in HCC tumours, correlate with SPHK1/2 expression and T cell dysfunction and associate with poor patient prognosis.

CONCLUSION: Our study uncovers the S1P-TREM2-NRF2 axis as a critical metabolic-immune circuit that preserves neutrophil survival and immunosuppressive function in HCC. Targeting this lipid-dependent ferroptosis resistance pathway offers a promising therapeutic strategy to overcome immunotherapy resistance in liver cancer.

PMID:42705697 | DOI:10.1136/gutjnl-2025-337414

S1P-TREM2 axis protects immunosuppressive neutrophils from ferroptosis to promote tumour progression in hepatocellular carcinoma

Gut. 2026 Sep 7:gutjnl-2025-337414. doi: 10.1136/gutjnl-2025-337414. Online ahead of print.

ABSTRACT

BACKGROUND: Neutrophils are increasingly recognised as immunosuppressive drivers of hepatocellular carcinoma (HCC), yet their persistence in the oxidative, lipid-rich tumour microenvironment remains poorly understood.

OBJECTIVE: To elucidate the metabolic and molecular programmes that enable tumour-associated neutrophils (TANs) to resist ferroptosis and sustain immunosuppression in HCC.

DESIGN: We employed human HCC samples, multiple murine HCC models, transcriptomic and lipidomic profiling, genetic loss-of-function systems and therapeutic interventions. Ferroptosis sensitivity, lipid metabolic rewiring and immunological consequences of TANs were systematically evaluated across models and validated in patient datasets and biospecimens.

RESULTS: TANs in human HCC and mouse models exhibit pronounced lipid accumulation and oxidative stress compared with peripheral neutrophils. Multi-omic profiling revealed that TANs are enriched for lipid-binding gene programmes and undergo rewiring towards sphingolipid and unsaturated fatty acid metabolism. We identified triggering receptor expressed on myeloid cells 2 (TREM2) as a key lipid-sensing receptor selectively expressed in TANs. Functional deletion of TREM2 reprogrammed the tumour immune microenvironment, restoring CD8+ T cell activity and suppressing HCC progression. Mechanistically, tumour-derived sphingosine-1-phosphate (S1P) activates TREM2, triggering nuclear factor erythroid 2-related factor 2 (NRF2)-mediated transcription of glutathione peroxidase 4 (GPX4) and solute carrier family 7 member 11 (SLC7A11), thereby promoting ferroptosis resistance. TREM2 expression is transcriptionally induced by granulocyte-macrophage colony-stimulating factor-signal transducer and activator of transcription 3 (GM-CSF-STAT3) signalling. Genetic deletion of TREM2, clustered regularly interspaced short palindromic repeats/CRISPR-associated protein 9 (CRISPR/Cas9)-mediated knockout of sphingosine kinase 1/2 (SPHK1/2) in tumour cells, or pharmacological inhibition of S1P synthesis disrupts this protective lipid-immune circuit, sensitises TANs to ferroptosis and restricts tumour growth. Therapeutically, a peptide-based TREM2 inhibitor reprogrammes TANs, restores CD8+ T cell function and enhances anti-programmed cell death protein 1 (PD-1) immunotherapy efficacy. Clinically, TREM2+ polymorphonuclear myeloid-derived suppressor cells (PMN-MDSCs) are enriched in HCC tumours, correlate with SPHK1/2 expression and T cell dysfunction and associate with poor patient prognosis.

CONCLUSION: Our study uncovers the S1P-TREM2-NRF2 axis as a critical metabolic-immune circuit that preserves neutrophil survival and immunosuppressive function in HCC. Targeting this lipid-dependent ferroptosis resistance pathway offers a promising therapeutic strategy to overcome immunotherapy resistance in liver cancer.

PMID:42705697 | DOI:10.1136/gutjnl-2025-337414

Nonlinear atomic tunnelling boosted by bright squeezed vacuum

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10485-9

Bright squeezed vacuum light boosts nonlinear atomic tunnelling ionization more than 20-fold compared with coherent light, enabling quantum control of strong-field processes without increasing classical intensity.

QED-Nano: Teaching a Tiny Model to Prove Hard Theorems

arXiv:2604.04898v1 Announce Type: new Abstract: Proprietary AI systems have recently demonstrated impressive capabilities on complex proof-based problems, with gold-level performance reported at the 2025 International Mathematical Olympiad (IMO). However, the training pipelines behind these systems remain largely undisclosed, and their reliance on large "internal" models and scaffolds makes them expensive to run, difficult to reproduce, and hard to study or improve upon. This raises a central question: can small, open models also be trained to achieve competitive reasoning performance on difficult Olympiad-level math? In this paper, we answer this question by building QED-Nano, a 4B model post-trained for Olympiad-level proofs. Our training recipe has three stages: (1) supervised fine-tuning to imbue good proof-writing styles by distilling from DeepSeek-Math-V2, (2) reinforcement learning (RL) with rubric-based rewards, and (3) expanding RL with a reasoning cache, which decomposes long proofs into iterative summarize-and-refine cycles and enables stronger test-time reasoning. QED-Nano surpasses the proof-generation performance of much larger open models, including Nomos-1 and GPT-OSS-120B, and approaches the performance of proprietary models like Gemini 3 Pro, at a fraction of the inference cost. To support further research on open mathematical reasoning, we release the full QED-Nano pipeline, including the QED-Nano and QED-Nano-SFT models, the FineProofs-SFT and FineProofs-RL datasets, and the training and evaluation code.

Stabilizing Unsupervised Self-Evolution of MLLMs via Continuous Softened Retracing reSampling

arXiv:2604.03647v1 Announce Type: cross Abstract: In the unsupervised self-evolution of Multimodal Large Language Models, the quality of feedback signals during post-training is pivotal for stable and effective learning. However, existing self-evolution methods predominantly rely on majority voting to select the most frequent output as the pseudo-golden answer, which may stem from the model's intrinsic biases rather than guaranteeing the objective correctness of the reasoning paths. To counteract the degradation, we propose \textbf{C}ontinuous \textbf{S}oftened \textbf{R}etracing re\textbf{S}ampling (\textbf{CSRS}) in MLLM self-evolution. Specifically, we introduce a Retracing Re-inference Mechanism (\textbf{RRM}) that the model re-inferences from anchor points to expand the exploration of long-tail reasoning paths. Simultaneously, we propose Softened Frequency Reward (\textbf{SFR}), which replaces binary rewards with continuous signals, calibrating reward based on the answers' frequency across sampled reasoning sets. Furthermore, incorporated with Visual Semantic Perturbation (\textbf{VSP}), CSRS ensures the model prioritizes mathematical logic over visual superficiality. Experimental results demonstrate that CSRS significantly enhances the reasoning performance of Qwen2.5-VL-7B on benchmarks such as MathVision. We achieve state-of-the-art (SOTA) results in unsupervised self-evolution on geometric tasks. Our code is avaible at https://github.com/yyy195/CSRS.

FURINA: A Fully Customizable Role-Playing Benchmark via Scalable Multi-Agent Collaboration Pipeline

arXiv:2510.06800v3 Announce Type: replace-cross Abstract: As large language models (LLMs) advance in role-playing (RP) tasks, existing benchmarks quickly become obsolete due to their narrow scope, outdated interaction paradigms, and limited adaptability across diverse application scenarios. To address this gap, we introduce FURINA-Builder, a novel multi-agent collaboration pipeline that automatically constructs fully customizable RP benchmarks at any scale. It enables evaluation of arbitrary characters across diverse scenarios and prompt formats, as the first benchmark builder in RP area for adaptable assessment. FURINA-Builder simulates dialogues between a test character and other characters drawn from a well-constructed character-scene pool, while an LLM judge selects fine-grained evaluation dimensions and adjusts the test character's responses into final test utterances. Using this pipeline, we build FURINA-Bench, a new comprehensive role-playing benchmark featuring both established and synthesized test characters, each assessed with dimension-specific evaluation criteria. Human evaluation and preliminary separability analysis justify our pipeline and benchmark design. We conduct extensive evaluations of cutting-edge LLMs and find that o3 and DeepSeek-R1 achieve the best performance on English and Chinese RP tasks, respectively. Across all models, established characters consistently outperform synthesized ones, with reasoning capabilities further amplifying this disparity. Interestingly, we observe that model scale does not monotonically reduce hallucinations. More critically, for reasoning LLMs, we uncover a novel trade-off: reasoning improves RP performance but simultaneously increases RP hallucinations. This trade-off extends to a broader Pareto frontier between RP performance and reliability for all LLMs. These findings demonstrate the effectiveness of FURINA-Builder and the challenge posed by FURINA-Bench.

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models

arXiv:2510.15148v2 Announce Type: replace-cross Abstract: Omni-modal large language models (OLLMs) aim to unify audio, vision, and text understanding within a single framework. While existing benchmarks primarily evaluate general cross-modal question-answering ability, it remains unclear whether OLLMs achieve modality-invariant reasoning or exhibit modality-specific biases. We introduce XModBench, a large-scale tri-modal benchmark explicitly designed to measure cross-modal consistency. XModBench comprises 60,828 multiple-choice questions spanning five task families and systematically covers all six modality compositions in question-answer pairs, enabling fine-grained diagnosis of an OLLM's modality-invariant reasoning, modality disparity, and directional imbalance. Experiments show that even the strongest model, Gemini 2.5 Pro, (i) struggles with spatial and temporal reasoning, achieving less than 60% accuracy, (ii) reveals persistent modality disparities, with performance dropping substantially when the same semantic content is conveyed through audio rather than text, and (iii) shows systematic directional imbalance, exhibiting lower consistency when vision serves as context compared to text. These findings indicate that current OLLMs remain far from truly modality-invariant reasoning and position XModBench as a fundamental diagnostic tool for evaluating and improving cross-modal competence. All data and evaluation tools will be available at https://xingruiwang.github.io/projects/XModBench/.

MOON3.0: Reasoning-aware Multimodal Representation Learning for E-commerce Product Understanding

arXiv:2604.00513v2 Announce Type: replace-cross Abstract: With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention. Although recent multimodal large language models (MLLMs) have driven significant progress in product understanding, they are typically employed as feature extractors that implicitly encode product information into global embeddings, thereby limiting their ability to capture fine-grained attributes. Therefore, we argue that leveraging the reasoning capabilities of MLLMs to explicitly model fine-grained product attributes holds significant potential. Nevertheless, achieving this goal remains non-trivial due to several key challenges: (i) long-context reasoning tends to dilute the model's attention to salient information in the raw input; (ii) supervised fine-tuning (SFT) primarily encourages rigid imitation, limiting the exploration of effective reasoning strategies; and (iii) fine-grained details are progressively attenuated during forward propagation. To address these issues, we propose MOON3.0, the first reasoning-aware MLLM-based model for product representation learning. Our method (1) employs a multi-head modality fusion module to adaptively integrate raw signals; (2) incorporates a joint contrastive and reinforcement learning framework to autonomously explore more effective reasoning strategies; and (3) introduces a fine-grained residual enhancement module to progressively preserve local details throughout the network. Additionally, we release a large-scale multimodal e-commerce benchmark MBE3.0. Experimentally, our model demonstrates state-of-the-art zero-shot performance across various downstream tasks on both our benchmark and public datasets.

Genetically encoded fluorescent reporters to visualize α-synuclein pathology in live brain

The development of genetically encoded fluorescent reporters, along with their corresponding knock-in mouse lines for labeling α-Syn inclusions, enables diverse applications in studying the propagation and pathological effects of α-Syn inclusions in the live brain.

Catgut implantation at acupoints improves anti-PD-1 inhibitor efficacy in lung cancer by inducing immune responses and remodeling the tumor microenvironment

Cancer Immunol Immunother. 2026 Mar 31;75(4):126. doi: 10.1007/s00262-026-04368-1.

ABSTRACT

While anti-programmed death-1 (anti-PD-1) therapy has revolutionized lung cancer treatment, its efficacy remains limited by an immunosuppressive tumor microenvironment (TME). We therefore investigated whether combining anti-PD-1 inhibitor with catgut embedding at the Zusanli acupoint (CIAA) could enhance anti-tumor immunity by reprogramming the TME in a lung cancer mouse model. Combining in vivo tumor monitoring, multi-parametric immune profiling (flow cytometry, IHC, ELISA), and multi-omics analyses (transcriptomics and metabolomics), we found that the combination therapy was associated with enhanced tumor growth inhibition. This effect correlated with a comprehensive TME transformation: conversion to an immunologically active state with increased effector immune cell infiltration (CD8⁺ T, CD4⁺ T, B cells, macrophages) and decreased regulatory T cells, coupled with suppression of pro-tumorigenic factors (VEGF, IL-6). Integrated omics analysis suggests that the combined treatment may modulate tumor-stroma interaction pathways (e.g., PI3K-Akt, focal adhesion) and rewire immunometabolic networks (e.g., tryptophan metabolism). Our study provides hypothesis-generating correlative data positioning CIAA as a potential adjunct capable of remodeling the TME to potentiate anti-PD-1 therapy in lung cancer.

PMID:41915222 | PMC:PMC13038699 | DOI:10.1007/s00262-026-04368-1

Catgut implantation at acupoints improves anti-PD-1 inhibitor efficacy in lung cancer by inducing immune responses and remodeling the tumor microenvironment

Cancer Immunol Immunother. 2026 Mar 31;75(4):126. doi: 10.1007/s00262-026-04368-1.

ABSTRACT

While anti-programmed death-1 (anti-PD-1) therapy has revolutionized lung cancer treatment, its efficacy remains limited by an immunosuppressive tumor microenvironment (TME). We therefore investigated whether combining anti-PD-1 inhibitor with catgut embedding at the Zusanli acupoint (CIAA) could enhance anti-tumor immunity by reprogramming the TME in a lung cancer mouse model. Combining in vivo tumor monitoring, multi-parametric immune profiling (flow cytometry, IHC, ELISA), and multi-omics analyses (transcriptomics and metabolomics), we found that the combination therapy was associated with enhanced tumor growth inhibition. This effect correlated with a comprehensive TME transformation: conversion to an immunologically active state with increased effector immune cell infiltration (CD8⁺ T, CD4⁺ T, B cells, macrophages) and decreased regulatory T cells, coupled with suppression of pro-tumorigenic factors (VEGF, IL-6). Integrated omics analysis suggests that the combined treatment may modulate tumor-stroma interaction pathways (e.g., PI3K-Akt, focal adhesion) and rewire immunometabolic networks (e.g., tryptophan metabolism). Our study provides hypothesis-generating correlative data positioning CIAA as a potential adjunct capable of remodeling the TME to potentiate anti-PD-1 therapy in lung cancer.

PMID:41915222 | DOI:10.1007/s00262-026-04368-1

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

arXiv:2511.12449v2 Announce Type: replace-cross Abstract: Recent Multimodal Large Language Models (MLLMs) have significantly advanced e-commerce product understanding. However, they still face three challenges: (i) the modality imbalance induced by modality mixed training; (ii) underutilization of the intrinsic alignment relationships among visual and textual information within a product; and (iii) limited handling of noise in e-commerce multimodal data. To address these, we propose MOON2.0, a dynamic modality-balanced MultimOdal representation learning framework for e-commerce prOduct uNderstanding. It comprises: (1) a Modality-driven Mixture-of-Experts (MoE) that adaptively processes input samples by their modality composition, enabling Multimodal Joint Learning to mitigate the modality imbalance; (2) a Dual-level Alignment method to better leverage semantic alignment properties inside individual products; and (3) an MLLM-based Image-text Co-augmentation strategy that integrates textual enrichment with visual expansion, coupled with Dynamic Sample Filtering to improve training data quality. We further release MBE2.0, a co-augmented Multimodal representation Benchmark for E-commerce representation learning and evaluation at https://huggingface.co/datasets/ZHNie/MBE2.0. Experiments show that MOON2.0 delivers state-of-the-art zero-shot performance on MBE2.0 and multiple public datasets. Furthermore, attention-based heatmap visualization provides qualitative evidence of improved multimodal alignment of MOON2.0.

MKA: Memory-Keyed Attention for Efficient Long-Context Reasoning

arXiv:2603.20586v2 Announce Type: replace-cross Abstract: As long-context language modeling becomes increasingly important, the cost of maintaining and attending to large Key/Value (KV) caches grows rapidly, becoming a major bottleneck in both training and inference. While prior works such as Multi-Query Attention (MQA) and Multi-Latent Attention (MLA) reduce memory by sharing or compressing KV features, they often trade off representation quality or incur runtime overhead. We propose Memory-Keyed Attention (MKA), a hierarchical attention mechanism that integrates multi-level KV caches (local, session, and long-term) and learns to route attention across them dynamically. We further introduce Route-Fused MKA (FastMKA), a broadcast-routed variant that fuses memory sources before attention computation for improved efficiency. Experiments on different sequence lengths show that FastMKA achieves a favorable accuracy-efficiency trade-off: comparable perplexity to MLA while achieving up to 5x faster training throughput and 1.8x lower evaluation latency. These results highlight MKA as a practical and extensible framework for efficient long-context attention.

When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning

arXiv:2603.21289v2 Announce Type: replace-cross Abstract: Recent progress in multimodal large language models has led to strong performance on reasoning tasks, but these improvements largely rely on high-quality annotated data or teacher-model distillation, both of which are costly and difficult to scale. To address this, we propose an unsupervised self-evolution training framework for multimodal reasoning that achieves stable performance improvements without using human-annotated answers or external reward models. For each input, we sample multiple reasoning trajectories and jointly model their within group structure. We use the Actor's self-consistency signal as a training prior, and introduce a bounded Judge based modulation to continuously reweight trajectories of different quality. We further model the modulated scores as a group level distribution and convert absolute scores into relative advantages within each group, enabling more robust policy updates. Trained with Group Relative Policy Optimization (GRPO) on unlabeled data, our method consistently improves reasoning performance and generalization on five mathematical reasoning benchmarks, offering a scalable path toward self-evolving multimodal models. The code are available at https://github.com/OPPO-Mente-Lab/LLM-Self-Judge.

GCGNet: Graph-Consistent Generative Network for Time Series Forecasting with Exogenous Variables

arXiv:2603.08032v1 Announce Type: cross Abstract: Exogenous variables offer valuable supplementary information for predicting future endogenous variables. Forecasting with exogenous variables needs to consider both past-to-future dependencies (i.e., temporal correlations) and the influence of exogenous variables on endogenous variables (i.e., channel correlations). This is pivotal when future exogenous variables are available, because they may directly affect the future endogenous variables. Many methods have been proposed for time series forecasting with exogenous variables, focusing on modeling temporal and channel correlations. However, most of them use a two-step strategy, modeling temporal and channel correlations separately, which limits their ability to capture joint correlations across time and channels. Furthermore, in real-world scenarios, time series are frequently affected by various forms of noises, underscoring the critical importance of robustness in such correlations modeling. To address these limitations, we propose GCGNet, a Graph-Consistent Generative Network for time series forecasting with exogenous variables. Specifically, GCGNet first employs a Variational Generator to produce coarse predictions. A Graph Structure Aligner then further guides it by evaluating the consistency between the generated and true correlations, where the correlations are represented as graphs, and are robust to noises. Finally, a Graph Refiner is proposed to refine the predictions to prevent degeneration and improve accuracy. Extensive experiments on 12 real-world datasets demonstrate that GCGNet outperforms state-of-the-art baselines.

Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks

arXiv:2510.19195v4 Announce Type: replace-cross Abstract: Recent advancements in driving world models enable controllable generation of high-quality RGB videos or multimodal videos. Existing methods primarily focus on metrics related to generation quality and controllability. However, they often overlook the evaluation of downstream perception tasks, which are $\mathbf{really\ crucial}$ for the performance of autonomous driving. Existing methods usually leverage a training strategy that first pretrains on synthetic data and finetunes on real data, resulting in twice the epochs compared to the baseline (real data only). When we double the epochs in the baseline, the benefit of synthetic data becomes negligible. To thoroughly demonstrate the benefit of synthetic data, we introduce Dream4Drive, a novel synthetic data generation framework designed for enhancing the downstream perception tasks. Dream4Drive first decomposes the input video into several 3D-aware guidance maps and subsequently renders the 3D assets onto these guidance maps. Finally, the driving world model is fine-tuned to produce the edited, multi-view photorealistic videos, which can be used to train the downstream perception models. Dream4Drive enables unprecedented flexibility in generating multi-view corner cases at scale, significantly boosting corner case perception in autonomous driving. To facilitate future research, we also contribute a large-scale 3D asset dataset named DriveObj3D, covering the typical categories in driving scenarios and enabling diverse 3D-aware video editing. We conduct comprehensive experiments to show that Dream4Drive can effectively boost the performance of downstream perception models under various training epochs. Page: https://wm-research.github.io/Dream4Drive/ GitHub Link: https://github.com/wm-research/Dream4Drive

Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy

arXiv:2510.08646v2 Announce Type: replace-cross Abstract: Safety alignment of large language models currently faces a central challenge: existing alignment techniques often prioritize mitigating responses to harmful prompts at the expense of overcautious behavior, leading models to incorrectly refuse benign requests. A key goal of safe alignment is therefore to improve safety while simultaneously minimizing false refusals. In this work, we introduce Energy Landscape Steering (ELS), a novel, fine-tuning free framework designed to resolve this challenge through dynamic, inference-time intervention. We train a lightweight external Energy-Based Model (EBM) to assign high energy to undesirable states (false refusal or jailbreak) and low energy to desirable states (helpful response or safe reject). During inference, the EBM maps the LLM's internal activations to an energy landscape, and we use the gradient of the energy function to steer the hidden states toward low-energy regions in real time. This dynamically guides the model toward desirable behavior without modifying its parameters. By decoupling behavioral control from the model's core knowledge, ELS provides a flexible and computationally efficient solution. Extensive experiments across diverse models demonstrate its effectiveness, raising compliance on the ORB-H benchmark from 57.3 percent to 82.6 percent while maintaining baseline safety performance. Our work establishes a promising paradigm for building LLMs that simultaneously achieve high safety and low false refusal rates.

Learning from Complexity: Exploring Dynamic Sample Pruning of Spatio-Temporal Training

arXiv:2602.19113v1 Announce Type: cross Abstract: Spatio-temporal forecasting is fundamental to intelligent systems in transportation, climate science, and urban planning. However, training deep learning models on the massive, often redundant, datasets from these domains presents a significant computational bottleneck. Existing solutions typically focus on optimizing model architectures or optimizers, while overlooking the inherent inefficiency of the training data itself. This conventional approach of iterating over the entire static dataset each epoch wastes considerable resources on easy-to-learn or repetitive samples. In this paper, we explore a novel training-efficiency techniques, namely learning from complexity with dynamic sample pruning, ST-Prune, for spatio-temporal forecasting. Through dynamic sample pruning, we aim to intelligently identify the most informative samples based on the model's real-time learning state, thereby accelerating convergence and improving training efficiency. Extensive experiments conducted on real-world spatio-temporal datasets show that ST-Prune significantly accelerates the training speed while maintaining or even improving the model performance, and it also has scalability and universality.

ST-EVO: Towards Generative Spatio-Temporal Evolution of Multi-Agent Communication Topologies

arXiv:2602.14681v2 Announce Type: replace-cross Abstract: LLM-powered Multi-Agent Systems (MAS) have emerged as an effective approach towards collaborative intelligence, and have attracted wide research interests. Among them, ``self-evolving'' MAS, treated as a more flexible and powerful technical route, can construct task-adaptive workflows or communication topologies, instead of relying on a predefined static structue template. Current self-evolving MAS mainly focus on Spatial Evolving or Temporal Evolving paradigm, which only considers the single dimension of evolution and does not fully incentivize LLMs' collaborative capability. In this work, we start from a novel Spatio-Temporal perspective by proposing ST-EVO, which supports dialogue-wise communication scheduling with a compact yet powerful flow-matching based Scheduler. To make precise Spatio-Temporal scheduling, ST-EVO can also perceive the uncertainty of MAS, and possesses self-feedback ability to learn from accumulated experience. Extensive experiments on nine benchmarks demonstrate the state-of-the-art performance of ST-EVO, achieving about 5%--25% accuracy improvement.
❌