❌

Normal view

RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems

arXiv:2609.09657v1 Announce Type: new Abstract: Existing emotional support conversation systems mainly focus on one-on-one seeker-supporter interactions and individual emotional states, leaving interpersonal relations in multi-party scenarios underexplored. In this work, we introduce relation-aware emotional support conversation, a new task that evaluates whether LLMs can capture and utilize the evolving dynamics of relationships to offer more effective emotional support. We construct RESCUE (Relation-aware Emotional Support Conversation Understanding and Evaluation Benchmark) from real couple and family interview conversations, containing 191 samples, 7,079 annotated turns, and 1,064.8 minutes of video. Based on rich annotations of socio-emotional and support-related dynamics, RESCUE defines six tasks that evaluate two core capabilities required for relation-aware emotional support: Relational Understanding and Relation-Sensitive Support. Experiments with ten LLMs show that current models perform relatively well on tasks relying on local emotional or intervention cues, but struggle with relation-intensive tasks such as relation pattern prediction, viewpoint prediction, and support strategy prediction. These findings reveal the limitations of current LLMs in modeling interpersonal relations and making relation-sensitive support decisions.

Early stage nonsmall cell lung cancer: Toward a risk-adaptive paradigm in the era of biologic precision

CA Cancer J Clin. 2026 Sep-Oct;76(5):e70100. doi: 10.3322/caac.70100.

ABSTRACT

The clinical landscape of early stage nonsmall cell lung cancer is at transformative crossroads. Driven by the widespread adoption of low-dose computed tomography screening, the frequent detection of ground-glass opacities, and a rising incidence among never-smokers, the diagnostic center of gravity has shifted toward earlier, potentially curable disease. This shift has been accompanied by equally important therapeutic advances, including parenchyma-sparing surgical techniques, minimally invasive platforms enhanced by digital navigation, and the transformative integration of perioperative immunotherapy and targeted agents. Concurrently, noninvasive monitoring approaches, such as liquid biopsy, have emerged as powerful tools to guide precision management. Despite this progress, substantial barriers to achieving a universal cure persist. Clinicians continue to face uncertainty in the management of ground-glass opacities, the anatomy-based TNM staging system fails to capture the biologic heterogeneity of early tumors, and global disparities in access to innovation remain unresolved. To address these challenges, the authors propose a shift toward a risk-adaptive management paradigm that harnesses artificial intelligence-driven analytics and multi-omics profiling to tailor treatment intensity according to each patient's biologic risk. Such an approach would enable appropriate escalation for high-risk individuals while permitting safe de-escalation for those at low risk. This holistic, lifespan-oriented strategy must be embraced to deliver equitable and durable cures for patients with early stage nonsmall cell lung cancer.

PMID:42713910 | PMC:PMC13555834 | DOI:10.3322/caac.70100

Early stage nonsmall cell lung cancer: Toward a risk-adaptive paradigm in the era of biologic precision

CA Cancer J Clin. 2026 Sep-Oct;76(5):e70100. doi: 10.3322/caac.70100.

ABSTRACT

The clinical landscape of early stage nonsmall cell lung cancer is at transformative crossroads. Driven by the widespread adoption of low-dose computed tomography screening, the frequent detection of ground-glass opacities, and a rising incidence among never-smokers, the diagnostic center of gravity has shifted toward earlier, potentially curable disease. This shift has been accompanied by equally important therapeutic advances, including parenchyma-sparing surgical techniques, minimally invasive platforms enhanced by digital navigation, and the transformative integration of perioperative immunotherapy and targeted agents. Concurrently, noninvasive monitoring approaches, such as liquid biopsy, have emerged as powerful tools to guide precision management. Despite this progress, substantial barriers to achieving a universal cure persist. Clinicians continue to face uncertainty in the management of ground-glass opacities, the anatomy-based TNM staging system fails to capture the biologic heterogeneity of early tumors, and global disparities in access to innovation remain unresolved. To address these challenges, the authors propose a shift toward a risk-adaptive management paradigm that harnesses artificial intelligence-driven analytics and multi-omics profiling to tailor treatment intensity according to each patient's biologic risk. Such an approach would enable appropriate escalation for high-risk individuals while permitting safe de-escalation for those at low risk. This holistic, lifespan-oriented strategy must be embraced to deliver equitable and durable cures for patients with early stage nonsmall cell lung cancer.

PMID:42713910 | DOI:10.3322/caac.70100

From the invasive front to organotropic pre-metastatic niches: spatial immune regulatory networks governing cholangiocarcinoma dissemination and metastasis-intercepting immunotherapy

4 September 2026 at 18:00

Front Immunol. 2026 Aug 20;17:1919864. doi: 10.3389/fimmu.2026.1919864. eCollection 2026.

ABSTRACT

Cholangiocarcinoma is an aggressive biliary tract malignancy in which metastatic relapse and primary or acquired resistance to immunotherapy remain major causes of mortality. Although immune checkpoint inhibitors have improved first-line treatment for advanced biliary tract cancer, most patients do not achieve durable benefit, indicating that immune failure is not explained by a single checkpoint pathway. In this Review, we propose a spatial immune-regulatory continuum for cholangiocarcinoma dissemination. Most direct single-cell and spatial evidence currently derives from intrahepatic cholangiocarcinoma, and its applicability to perihilar and distal disease remains to be established. This continuum begins in the tumor core and invasive front, where malignant cells, cancer-associated fibroblasts, tumor-associated macrophages, endothelial and lymphatic cells, regulatory T cells, immature neutrophils and excluded or dysfunctional cytotoxic T cells form a pro-invasive ecosystem. It then extends through extracellular vesicles, soluble mediators and lymphovascular routes that may educate organotropic pre-metastatic niches. Finally, lymph node, lung, liver, peritoneal and bone microenvironments provide organ-specific extracellular matrix, myeloid and stromal programs that enable immune evasion and metastatic colonization. By integrating clinical evidence, multi-omics studies, single-cell and spatial transcriptomics, extracellular vesicle biology, pre-metastatic niche concepts and emerging therapeutic strategies, we argue that cholangiocarcinoma metastasis should be targeted before overt dissemination whenever possible. In this Review, "metastasis-intercepting immunotherapy" is used as an author-defined conceptual framework for strategies intended to prevent or disrupt the immune-stromal conditions that enable dissemination and colonization, rather than merely shrink established metastatic lesions. Metastasis-intercepting immunotherapy will likely require rational combinations that reprogram the invasive front, restore dendritic-cell-mediated antigen presentation, block tumor-stroma-myeloid circuits, disrupt EV-mediated communication that may contribute to niche formation and select patients using spatial biomarkers rather than bulk immune markers alone.

PMID:42694469 | PMC:PMC13539491 | DOI:10.3389/fimmu.2026.1919864

Nitrogen dioxide exposure promotes CD8(+)T cell infiltration and contributes to increased susceptibility to ulcerative colitis: An integrative multi-omics, artificial intelligence, and mouse model study

J Hazard Mater. 2026 Sep 15;516:143449. doi: 10.1016/j.jhazmat.2026.143449. Epub 2026 Aug 30.

ABSTRACT

The global incidence of ulcerative colitis (UC) has significantly increased in rapidly industrializing nations, with numerous studies highlighting environmental exposures, particularly nitrogen dioxide (NO2), as potential contributors to disease susceptibility. However, the clinical implications and molecular mechanisms linking NO2 exposure to UC susceptibility remain poorly understood. This study investigated the associations between NO2 and UC by integrating multi-omics data. We identified a CD8+ T cell subpopulation with a distinct phenotype characterized by perforin production, which potentially exacerbated colonic inflammation related to NO2 exposure. To validate this hypothesis, we established mouse models exposed to NO2, confirming increased CD8+ T cell infiltration and elevated perforin secretion through immunofluorescent (IF) staining. Employing artificial intelligence techniques, we identified Cell Division Cycle 25B (CDC25B) as a gene of interest correlated with putative NO2-related UC signatures. Finally, through molecular docking (MD) and molecular dynamics simulations (MDS), we identified ozanimod as one of several computationally nominated compounds associated with the CDC25B‑related network; however, none of these computational predictions were experimentally validated in the present study. Collectively, these findings suggest a correlative link between perforin or CD8+ T cell-associated colonic inflammation and NO2-associated UC susceptibility, and nominate CDC25B as a candidate gene for further investigation.

PMID:42679583 | DOI:10.1016/j.jhazmat.2026.143449

Dynamic Dual-Granularity Skill Bank for Agentic RL

arXiv:2603.28716v2 Announce Type: replace Abstract: Agentic RL can benefit substantially from reusable experience, yet existing skill-based methods mainly extract trajectory-level guidance and often lack principled mechanisms for maintaining an evolving skill memory. We propose D2Skill, a dynamic dual-granularity skill bank for agentic RL that organizes reusable experience into task skills for high-level guidance and step skills for fine-grained decision support and error correction. D2Skill jointly trains the policy and skill bank through paired baseline and skill-injected rollouts under the same policy, using their performance gap to derive hindsight utility signals for both skill updating and policy optimization. Built entirely from training-time experience, the skill bank is continuously expanded through reflection and maintained with utility-aware retrieval and pruning. Experiments on ALFWorld, WebShop, and Search-Augmented QA tasks show that D2Skill substantially improves performance over skill-free baselines across models of different scales. Further ablations and analyses show that both dual-granularity skill modeling and dynamic skill maintenance are critical to these gains, while the learned skills exhibit higher utility, transfer across evaluation settings, and introduce only modest training overhead.

SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment

arXiv:2604.08988v3 Announce Type: replace Abstract: Current LLM-based agents demonstrate strong performance in episodic task execution but remain constrained by static toolsets and episodic amnesia, failing to accumulate experience across task boundaries. This paper formalizes the Self-Evolving Agent (SEA) from the perspective of digital embodiment and continuous cross-task evolution, introduces the Evolutionary Flywheel as its minimal sufficient architecture, and presents SEA-Eval -- the first benchmark designed specifically for evaluating SEAs. Grounded in Flywheel theory, SEA-Eval establishes SR and T as primary metrics and, through sequential task stream design, is designed to quantify evolutionary gain, evolutionary stability, and implicit alignment convergence. Empirical evaluation reveals that, under comparable success rates, token consumption differs by up to 31.2 times between frameworks on individual tasks, with divergent evolutionary trajectories emerging under sequential analysis -- demonstrating that success rate alone creates a capability illusion and that the sequential convergence of $T$ is the key criterion for distinguishing genuine evolution from pseudo-evolution.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

arXiv:2602.18527v2 Announce Type: replace-cross Abstract: Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. This design choice introduces a fundamental dimensionality mismatch that precludes reliable source localization and spatial reasoning in complex 3D environments. We address this limitation by presenting JAEGER, a framework that extends AV-LLMs to 3D space, to enable joint spatial grounding and reasoning through the integration of RGB-D observations and multi-channel first-order ambisonics. A core contribution of our work is the neural intensity vector (Neural IV), a learned spatial audio representation that encodes robust directional cues to enhance direction-of-arrival estimation, even in adverse acoustic scenarios with overlapping sources. To facilitate large-scale training and systematic evaluation, we propose SpatialSceneQA, a benchmark of 61k instruction-tuning samples curated from simulated physical environments. Extensive experiments demonstrate that our approach consistently surpasses 2D-centric baselines across diverse spatial perception and reasoning tasks, underscoring the necessity of explicit 3D modelling for advancing AI in physical environments. Our source code, pre-trained model checkpoints, and datasets are available at https://github.com/liuzhan22/JAEGER.

Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning

arXiv:2605.12906v2 Announce Type: replace-cross Abstract: Data selection during supervised fine-tuning (SFT) can critically change the behavior of large language models (LLMs). Although existing work has studied the effect of selecting data based on heuristics such as perplexity, difficulty, or length, the reported findings are often inconsistent or context-dependent. In this work, we systematically study the role of data difficulty in fine-tuning from both empirical and theoretical perspectives, and find that there is no universally optimal difficulty level; rather, its effectiveness depends on the dataset size. We show that for a fixed data budget, there exists an optimal data difficulty for SFT, and that this optimal difficulty shifts toward harder data as the data budget increases. To explain this phenomenon, we conduct controlled synthetic experiments that reveal a simple underlying mechanism: the interplay between the (in-distribution) generalization gap and the extrapolation gap. We further support this mechanism through a theoretical analysis using PAC-Bayesian generalization bounds. Overall, our results clarify how data size and difficulty jointly affect the trade-off between generalization and extrapolation in SFT, providing guidance for difficulty-based data selection under certain model and data conditions.

FDX1 as a predictive biomarker and therapeutic target for lymph node metastasis in gastric cancer

Clin Exp Med. 2026 May 10. doi: 10.1007/s10238-026-02160-0. Online ahead of print.

ABSTRACT

The prognostic values of cuproptosis-related genes (CRGs) in gastric cancer with lymph node metastasis (GCLM), especially in the tumor immune microenvironment (TIME), remain unclear. We analyzed the expression, mutation, immunity, drug sensitivity, and prognostic value of CRGs in GCLM using TCGA and GEO cohorts. Consensus clustering was performed to identify CRG subtypes, with differences characterized by multi-omics analysis. A CRG-based prognostic risk score and immune score were constructed for individualized assessment, and the role of CRGs was validated through in vitro and in vivo experiments. Consensus clustering revealed that CRGs were significantly enriched in biological processes related to mitosis and energy metabolism, as well as in immune-related and cancer-associated pathways. Four distinct CRG subtypes were identified, showing marked differences in expression profiles, prognosis, genetic alterations, TIME, and chemotherapeutic drug sensitivity. We developed an exploratory CRG-based prognostic risk score for preliminary individualized assessment, and the functional relevance of CRGs in GCLM was further validated through in vitro experiments. Among these, FDX1, LIAS, DLAT, MTF1, and GLS were identified as key determinants of overall survival in patients with GCLM, with FDX1 emerging as a potential independent prognostic factor. Notably, upregulation of FDX1 significantly suppressed lymph node metastasis of gastric cancer cells in a mouse popliteal lymph node metastasis model. Our data uncovers FDX1 might be a potential favorable prognostic factors in GCLM patients. These findings may improve our understanding of CRGs in GCLM and provide new in-sights for assessing prognosis and developing more effective treatment strategies.

PMID:42107026 | DOI:10.1007/s10238-026-02160-0

ActionNex: A Virtual Outage Manager for Cloud

arXiv:2604.03512v1 Announce Type: new Abstract: Outage management in large-scale cloud operations remains heavily manual, requiring rapid triage, cross-team coordination, and experience-driven decisions under partial observability. We present \textbf{ActionNex}, a production-grade agentic system that supports end-to-end outage assistance, including real-time updates, knowledge distillation, and role- and stage-conditioned next-best action recommendations. ActionNex ingests multimodal operational signals (e.g., outage content, telemetry, and human communications) and compresses them into critical events that represent meaningful state transitions. It couples this perception layer with a hierarchical memory subsystem: long-term Key-Condition-Action (KCA) knowledge distilled from playbooks and historical executions, episodic memory of prior outages, and working memory of the live context. A reasoning agent aligns current critical events to preconditions, retrieves relevant memories, and generates actionable recommendations; executed human actions serve as an implicit feedback signal to enable continual self-evolution in a human-agent hybrid system. We evaluate ActionNex on eight real Azure outages (8M tokens, 4,000 critical events) using two complementary ground-truth action sets, achieving 71.4\% precision and 52.8-54.8\% recall. The system has been piloted in production and has received positive early feedback.

Schema-Aware Planning and Hybrid Knowledge Toolset for Reliable Knowledge Graph Triple Verification

arXiv:2604.04190v1 Announce Type: new Abstract: Knowledge Graphs (KGs) serve as a critical foundation for AI systems, yet their automated construction inevitably introduces noise, compromising data trustworthiness. Existing triple verification methods, based on graph embeddings or language models, often suffer from single-source bias by relying on either internal structural constraints or external semantic evidence, and usually follow a static inference paradigm. As a result, they struggle with complex or long-tail facts and provide limited interpretability. To address these limitations, we propose SHARP (Schema-Hybrid Agent for Reliable Prediction), a training-free autonomous agent that reformulates triple verification as a dynamic process of strategic planning, active investigation, and evidential reasoning. Specifically, SHARP combines a Memory-Augmented Mechanism with Schema-Aware Strategic Planning to improve reasoning stability, and employs an enhanced ReAct loop with a Hybrid Knowledge Toolset to dynamically integrate internal KG structure and external textual evidence for cross-verification. Experiments on FB15K-237 and Wikidata5M-Ind show that SHARP significantly outperforms existing state-of-the-art baselines, achieving accuracy gains of 4.2% and 12.9%, respectively. Moreover, SHARP provides transparent, fact-based evidence chains for each judgment, demonstrating strong interpretability and robustness for complex verification tasks.

Diagonal-Tiled Mixed-Precision Attention for Efficient Low-Bit MXFP Inference

arXiv:2604.03950v1 Announce Type: cross Abstract: Transformer-based large language models (LLMs) have demonstrated remarkable performance across a wide range of real-world tasks, but their inference cost remains prohibitively high due to the quadratic complexity of attention and the memory bandwidth limitations of high-precision operations. In this work, we present a low-bit mixed-precision attention kernel using the microscaling floating-point (MXFP) data format, utilizing the computing capability on next-generation GPU architectures. Our Diagonal-Tiled Mixed-Precision Attention (DMA) incorporates two kinds of low-bit computation at the tiling-level, and is a delicate fused kernel implemented using Triton, exploiting hardware-level parallelism and memory efficiency to enable fast and efficient inference without compromising model performance. Extensive empirical evaluations on NVIDIA B200 GPUs show that our kernel maintains generation quality with negligible degradation, and meanwhile achieves significant speedup by kernel fusion. We release our code at https://github.com/yifu-ding/MP-Sparse-Attn.

ROSClaw: A Hierarchical Semantic-Physical Framework for Heterogeneous Multi-Agent Collaboration

arXiv:2604.04664v1 Announce Type: cross Abstract: The integration of large language models (LLMs) with embodied agents has improved high-level reasoning capabilities; however, a critical gap remains between semantic understanding and physical execution. While vision-language-action (VLA) and vision-language-navigation (VLN) systems enable robots to perform manipulation and navigation tasks from natural language instructions, they still struggle with long-horizon sequential and temporally structured tasks. Existing frameworks typically adopt modular pipelines for data collection, skill training, and policy deployment, resulting in high costs in experimental validation and policy optimization. To address these limitations, we propose ROSClaw, an agent framework for heterogeneous robots that integrates policy learning and task execution within a unified vision-language model (VLM) controller. The framework leverages e-URDF representations of heterogeneous robots as physical constraints to construct a sim-to-real topological mapping, enabling real-time access to the physical states of both simulated and real-world agents. We further incorporate a data collection and state accumulation mechanism that stores robot states, multimodal observations, and execution trajectories during real-world execution, enabling subsequent iterative policy optimization. During deployment, a unified agent maintains semantic continuity between reasoning and execution, and dynamically assigns task-specific control to different agents, thereby improving robustness in multi-policy execution. By establishing an autonomous closed-loop framework, ROSClaw minimizes the reliance on robot-specific development workflows. The framework supports hardware-level validation, automated generation of SDK-level control programs, and tool-based execution, enabling rapid cross-platform transfer and continual improvement of robotic skills. Ours project page: https://www.rosclaw.io/.

Mind Your HEARTBEAT! Claw Background Execution Inherently Enables Silent Memory Pollution

arXiv:2603.23064v3 Announce Type: replace-cross Abstract: We identify a critical security vulnerability in mainstream Claw personal AI agents: untrusted content encountered during heartbeat-driven background execution can silently pollute agent memory and subsequently influence user-facing behavior without the user's awareness. This vulnerability arises from an architectural design shared across the Claw ecosystem: heartbeat background execution runs in the same session as user-facing conversation, so content ingested from any external source monitored in the background (including email, message channels, news feeds, code repositories, and social platforms) can enter the same memory context used for foreground interaction, often with limited user visibility and without clear source provenance. We formalize this process as an Exposure (E) $\rightarrow$ Memory (M) $\rightarrow$ Behavior (B) pathway: misinformation encountered during heartbeat execution enters the agent's short-term session context, potentially gets written into long-term memory, and later shapes downstream user-facing behavior. We instantiate this pathway in an agent-native social setting using MissClaw, a controlled research replica of Moltbook. We find that (1) social credibility cues, especially perceived consensus, are the dominant driver of short-term behavioral influence, with misleading rates up to 61%; (2) routine memory-saving behavior can promote short-term pollution into durable long-term memory at rates up to 91%, with cross-session behavioral influence reaching 76%; (3) under naturalistic browsing with content dilution and context pruning, pollution still crosses session boundaries. Overall, prompt injection is not required: ordinary social misinformation is sufficient to silently shape agent memory and behavior under heartbeat-driven background execution.

Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation

arXiv:2604.02324v1 Announce Type: cross Abstract: Language models (LMs) are increasingly extended with new learnable vocabulary tokens for domain-specific tasks, such as Semantic-ID tokens in generative recommendation. The standard practice initializes these new tokens as the mean of existing vocabulary embeddings, then relies on supervised fine-tuning to learn their representations. We present a systematic analysis of this strategy: through spectral and geometric diagnostics, we show that mean initialization collapses all new tokens into a degenerate subspace, erasing inter-token distinctions that subsequent fine-tuning struggles to fully recover. These findings suggest that \emph{token initialization} is a key bottleneck when extending LMs with new vocabularies. Motivated by this diagnosis, we propose the \emph{Grounded Token Initialization Hypothesis}: linguistically grounding novel tokens in the pretrained embedding space before fine-tuning better enables the model to leverage its general-purpose knowledge for novel-token domains. We operationalize this hypothesis as GTI (Grounded Token Initialization), a lightweight grounding stage that, prior to fine-tuning, maps new tokens to distinct, semantically meaningful locations in the pretrained embedding space using only paired linguistic supervision. Despite its simplicity, GTI outperforms both mean initialization and existing auxiliary-task adaptation methods in the majority of evaluation settings across multiple generative recommendation benchmarks, including industry-scale and public datasets. Further analyses show that grounded embeddings produce richer inter-token structure that persists through fine-tuning, corroborating the hypothesis that initialization quality is a key bottleneck in vocabulary extension.

ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation

arXiv:2603.29902v1 Announce Type: new Abstract: Interleaved text-and-image generation represents a significant frontier for Multimodal Large Language Models (MLLMs), offering a more intuitive way to convey complex information. Current paradigms rely on either image generation or retrieval augmentation, yet they typically treat the two as mutually exclusive paths, failing to unify factuality with creativity. We argue that the next milestone in this field is Agentic Tool Planning, where the model serves as a central controller that autonomously determines when, where, and which tools to invoke to produce interleaved responses for visual-critical queries. To systematically evaluate this paradigm, we introduce ATP-Bench, a novel benchmark comprising 7,702 QA pairs (including 1,592 VQA pairs) across eight categories and 25 visual-critical intents, featuring human-verified queries and ground truths. Furthermore, to evaluate agentic planning independent of end-to-end execution and changing tool backends, we propose a Multi-Agent MLLM-as-a-Judge (MAM) system. MAM evaluates tool-call precision, identifies missed opportunities for tool use, and assesses overall response quality without requiring ground-truth references. Our extensive experiments on 10 state-of-the-art MLLMs reveal that models struggle with coherent interleaved planning and exhibit significant variations in tool-use behavior, highlighting substantial room for improvement and providing actionable guidance for advancing interleaved generation. Dataset and code are available at https://github.com/Qwen-Applications/ATP-Bench.

MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines

arXiv:2603.06679v2 Announce Type: replace Abstract: Video world models have shown immense promise for interactive simulation and entertainment, but current systems still struggle with two important aspects of interactivity: user control over the environment for reproducible, editable experiences, and shared inference where players hold influence over a common world. To address these limitations, we introduce an explicit external memory into the system, a persistent state operating independent of the model's context window, that is continually updated by user actions and queried throughout the generation roll-out. Unlike conventional diffusion game engines that operate as next-frame predictors, our approach decomposes generation into Memory, Observation, and Dynamics modules. This design gives users direct, editable control over environment structure via an editable memory representation, and it naturally extends to real-time multiplayer rollouts with coherent viewpoints and consistent cross-player interactions.

QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentation

arXiv:2507.13266v4 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as a central paradigm for training large language models (LLMs) in reasoning tasks. Yet recent studies question RL's ability to incentivize reasoning capacity beyond the base model. This raises a key challenge: how can RL be adapted to solve harder reasoning problems more effectively? To address this challenge, we propose a simple yet effective strategy via Question Augmentation: introduce partial solutions during training to reduce problem difficulty and provide more informative learning signals. Our method, QuestA, when applied during RL training on math reasoning tasks, not only improves pass@1 but also pass@k-particularly on problems where standard RL struggles to make progress. This enables continual improvement over strong open-source models such as DeepScaleR and OpenMath Nemotron, further enhancing their reasoning capabilities. We achieve new state-of-the-art results on math benchmarks using 1.5B-parameter models: 72.50% (+10.73%) on AIME24, 62.29% (+12.79%) on AIME25, and 41.67% (+10.11%) on HMMT25. Code, data and model are available at https://github.com/foreverlasting1202/QuestA.
❌