❌

Reading view

Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

arXiv:2602.08499v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is an effective paradigm for improving the reasoning capabilities of large language models. However, existing RLVR methods utilize rollouts in an indiscriminate and short-horizon manner: responses of heterogeneous quality within each prompt are treated uniformly, and historical rollouts are discarded after a single use. This leads to noisy supervision, poor sample efficiency, and suboptimal policy updates. We address these issues by formulating rollout scheduling in RLVR as a contextual bandit problem and proposing a unified neural scheduling framework that adaptively selects high-value rollouts throughout training. Each rollout is treated as an arm whose reward is defined by the induced performance gain between consecutive optimization steps. The resulting scheduler supports both noise-aware intra-group selection and adaptive global reuse of historical rollouts within a single principled framework. We provide theoretical justification by deriving sublinear regret bounds and showing that enlarging the rollout buffer improves the achievable performance upper bound. Experiments on six mathematical reasoning benchmarks demonstrate consistent gains in performance and training efficiency across multiple RLVR optimization methods.
  •  

Low-dose intestinal irradiation enhances the efficacy and prognosis of PD-1 blockade in metastatic non-small cell lung cancer

Clin Cancer Res. 2026 Mar 18. doi: 10.1158/1078-0432.CCR-25-4153. Online ahead of print.

ABSTRACT

PURPOSE: Intestinal low-dose irradiation (ILDR) may enhance immunotherapy efficacy by modulating the gut microbiota and metabolism; however, its role in metastatic non-small cell lung cancer (mNSCLC), particularly in the first-line setting, remains unclear.

EXPERIMENTAL DESIGN: This multicenter retrospective and prospective study included mNSCLC patients receiving first- and second-line programmed cell death protein 1 (PD-1) inhibitors along with abdominopelvic radiotherapy between 2018 and 2025. Patients were stratified by the mean intestinal radiation dose into <1 Gy, 1-3 Gy, and >3 Gy groups and treatment outcomes were compared. The blood and fecal samples were subjected to multi-omics profiling.

RESULTS: g>309 patients were included in the retrospective analysis. Optimal efficacy was observed with a small intestinal mean radiation dose (SIMRD) of 1-3 Gy, showing longer progression-free survival (PFS, 10.2 months) and overall survival (OS, 22.8 months) (P < 0.01), which was consistent across subgroups. Compared with 1-3 Gy, SIMRD >3 Gy (Hazard ratio [HR] = 4.87, P < 0.001) and <1 Gy (HR = 1.85, P < 0.001) independently predicted worse OS. Prospective results confirmed the best disease control rate (P = 0.041) and PFS (P = 0.046) with SIMRD of 1-3 Gy. Responders were enriched in Bacillota, Clostridia, and indole derivatives, particularly indole-3-carboxylic acid. Moreover, the 1-3 Gy group exhibited increased circulating macrophage inflammatory protein-3Ξ± and reduced circulating Ξ±4Ξ²7+ regulatory T cells.

CONCLUSIONS: ILDR influences the efficacy of PD-1 blockade in patients with mNSCLC, particularly when SIMRD is maintained within the 1-3 Gy range, likely through modulation of the gut microbiota-metabolite-immune axis.

PMID:41849236 | DOI:10.1158/1078-0432.CCR-25-4153

  •  

CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling

arXiv:2603.08035v1 Announce Type: new Abstract: Reward modeling is essential for aligning Large Language Models(LLMs) with human preferences, yet conventional reward models suffer from poor interpretability and heavy reliance on costly expert annotations. While recent rubric-based approaches enhance evaluation transparency, they lack systematic quality control, yielding noisy and redundant criteria, failing to mitigate persistent biases (e.g., verbosity, position) in LLM evaluators, and creating a scalability-reliability trade-off. To address these limitations, we propose CDRRM (Contrast-Driven Rubric Reward Model), a framework built on a novel Contrast-then-Synthesis paradigm for high-quality rubric generation and guided preference judgment. CDRRM first conducts multi-dimensional contrastive profiling on preference pairs to identify causal discriminative factors, then synthesizes these insights into compact, context-aware rubrics to guide preference judg- ments. Extensive experiments on three authoritative benchmarks (RewardBench, RMBench, RMB) demonstrate that CDRRM achieves state-of-the-art performance across diverse domains and effectively mitigates aforementioned evaluation biases. Notably, our approach delivers exceptional data efficiency: training the rubric generator on only 3k high-quality samples empowers a frozen pre-trained judge model to outperform fully fine-tuned baselines. This work offers a scalable, interpretable, and data-efficient path for reward modeling.
  •  
❌