❌

Normal view

Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

arXiv:2602.08499v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is an effective paradigm for improving the reasoning capabilities of large language models. However, existing RLVR methods utilize rollouts in an indiscriminate and short-horizon manner: responses of heterogeneous quality within each prompt are treated uniformly, and historical rollouts are discarded after a single use. This leads to noisy supervision, poor sample efficiency, and suboptimal policy updates. We address these issues by formulating rollout scheduling in RLVR as a contextual bandit problem and proposing a unified neural scheduling framework that adaptively selects high-value rollouts throughout training. Each rollout is treated as an arm whose reward is defined by the induced performance gain between consecutive optimization steps. The resulting scheduler supports both noise-aware intra-group selection and adaptive global reuse of historical rollouts within a single principled framework. We provide theoretical justification by deriving sublinear regret bounds and showing that enlarging the rollout buffer improves the achievable performance upper bound. Experiments on six mathematical reasoning benchmarks demonstrate consistent gains in performance and training efficiency across multiple RLVR optimization methods.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

arXiv:2602.18527v2 Announce Type: replace-cross Abstract: Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. This design choice introduces a fundamental dimensionality mismatch that precludes reliable source localization and spatial reasoning in complex 3D environments. We address this limitation by presenting JAEGER, a framework that extends AV-LLMs to 3D space, to enable joint spatial grounding and reasoning through the integration of RGB-D observations and multi-channel first-order ambisonics. A core contribution of our work is the neural intensity vector (Neural IV), a learned spatial audio representation that encodes robust directional cues to enhance direction-of-arrival estimation, even in adverse acoustic scenarios with overlapping sources. To facilitate large-scale training and systematic evaluation, we propose SpatialSceneQA, a benchmark of 61k instruction-tuning samples curated from simulated physical environments. Extensive experiments demonstrate that our approach consistently surpasses 2D-centric baselines across diverse spatial perception and reasoning tasks, underscoring the necessity of explicit 3D modelling for advancing AI in physical environments. Our source code, pre-trained model checkpoints, and datasets are available at https://github.com/liuzhan22/JAEGER.

L-Drive: Beyond a Single Mapping-Latent Context Drives Time Series Forecasting

arXiv:2605.17730v2 Announce Type: replace-cross Abstract: Mainstream methods for multivariate time-series forecasting largely follow the Direct-Mapping paradigm. They learn a unified mapping from history to the future in the observation space to fit value-level dependencies. However, real-world systems often undergo distribution shifts and regime changes. In such cases, a unified mapping can exhibit response lag around turning points, causing error accumulation within the switching window and reducing forecasting reliability. To address this issue, we propose L-Drive, a change-aware forecasting framework. L-Drive introduces a Latent-Context, to explicitly characterize high-level dynamics evolving over time, and uses gating to modulate increment representations. This provides more timely change cues and improves adaptation to changing segments. In addition, it incorporates patch-shared relative positional basis functions to strengthen intra-segment structural modeling and reduce overfitting caused by absolute-position memorization. Extensive experiments validate the effectiveness of L-Drive and show a better overall trade-off between forecasting accuracy and computational efficiency.

Multi-omics Analysis Reveals the Correlation of Gut Microbiota and Metabolites With Thalidomide Treatment for Chemotherapy-Induced Nausea and Vomiting in Small Cell Lung Cancer

17 April 2026 at 18:00

Biotechnol J. 2026 Apr;21(4):e70228. doi: 10.1002/biot.70228.

ABSTRACT

Small cell lung cancer (SCLC) is a highly aggressive malignancy, and chemotherapy frequently causes nausea and vomiting, which can impair treatment tolerance. Because thalidomide (THD) has shown potential clinical benefit in alleviating nausea and anorexia, we investigated whether its effects might be associated with changes in gut microbial composition and metabolite profiles. Fecal samples were collected from patients with SCLC and categorized into THD-treated and control groups. Metagenomic sequencing and nontargeted metabolomic profiling were performed to characterize microbial composition and metabolic signatures. THD treatment was also associated with higher microbial alpha diversity and increased abundance of genera such as Eubacterium and Prevotella. Metabolomic analysis identified several differential metabolites, including hydrogenated MDI, becocalcidiol, β-octylglucoside, and azelaic acid. Collectively, these findings suggest that the gut microbiota-metabolite axis may be associated with the potential effects of THD on CINV and anorexia in patients with SCLC. The identified microbial taxa and metabolites may serve as candidate biomarkers or potential therapeutic targets, although further validation in larger studies is necessary.

PMID:41994961 | PMC:PMC13088213 | DOI:10.1002/biot.70228

❌