❌

Normal view

Self-supervised Hierarchical Visual Reasoning with World Model

arXiv:2605.17537v2 Announce Type: replace Abstract: 3D open-world environments with adversarial opponents remain a core challenge for reinforcement learning due to their vast state spaces. Effective reasoning representations are essential in such settings. While existing self-supervised visual foresight reasoning approaches often suffer from multi-step error accumulation, many recent studies resort to injecting domain-specific knowledge for more stable guidance. Our key insight is that the photorealistic fidelity of visual reasoning representations is secondary; what truly matters is providing informative, task-relevant signals. To this end, we propose ResDreamer, a hierarchical world model in which each higher-level layer is trained to reconstruct the residuals of the layer below. This design enables progressive abstraction of increasingly sophisticated world dynamics and fosters the emergence of richer latent representations. Drawing inspiration from the "Bitter Lesson", ResDreamer trains its reasoning representations in a purely self-supervised manner. The higher-level residual representations are used to modulate lower-level predictions, allowing the world model to scale effectively with only linearly increasing cross-layer communication costs. Experiments show that ResDreamer achieves state-of-the-art sample efficiency and parameter efficiency. This scalable hierarchical visual foresight reasoning architecture paves the way for more capable online RL agents in open-ended, dynamic environments. The code is accessible at https://github.com/XuYuanFei01/ResDreamer.

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens

arXiv:2602.15620v5 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) has significantly improved large language model reasoning, but existing RL fine-tuning methods rely heavily on heuristic techniques such as entropy regularization and reweighting to maintain stability. In practice, they often suffer from late-stage performance collapse, leading to degraded reasoning quality and unstable training. We identify a key factor behind this instability: a small fraction of tokens, termed spurious tokens (around 0.01%), which contribute little to the reasoning outcome but receive disproportionately amplified gradient updates due to inheriting the full sequence-level reward. We present a unified framework for evaluating token-level optimization impacts across spurious risk, gradient norms, and entropy changes. Building on the analysis of token characteristics that severely disrupt optimization, we propose the Silencing Spurious Tokens (S2T) mechanism to efficiently suppress their gradient perturbations. Incorporating this mechanism into a group-based objective, we propose Spurious-Token-Aware Policy Optimization (STAPO), which promotes stable and effective large-scale model refinement. Across six mathematical reasoning benchmarks using Qwen 1.7B, 8B, and 14B base models, STAPO consistently demonstrates superior entropy stability and achieves an average performance improvement of 11.49% ($\rho_{\mathrm{T}}$=1.0, top-p=1.0) and 3.73% ($\rho_{\mathrm{T}}$=0.7, top-p=0.9) over GRPO, 20-Entropy, and JustRL.

Circulating Tumor Cells in Pancreatic Ductal Adenocarcinoma: The Systemic Execution Hub of Metastasis

18 May 2026 at 18:00

Pharmacol Res. 2026 May 17:108253. doi: 10.1016/j.phrs.2026.108253. Online ahead of print.

ABSTRACT

Pancreatic ductal adenocarcinoma (PDAC) exemplifies early systemic dissemination, with circulating tumor cells (CTCs) at its core. We advance a unified conceptual framework that positions CTCs as the systemic execution hub of PDAC metastasis, dynamic entity that coordinates the metastatic cascade via four cardinal functions: Seeding, Adapting, Engineering, and Signaling. Integrating eco-evolutionary dynamics, this hub actively drives phenotypic selection, niche remodeling, and immune evasion, while providing real-time biologic intelligence through liquid biopsy. Robust clinical correlation has not yet translated into routine practice because of technical variability, biological complexity, and a lack of interventional evidence. We therefore propose an evidence-driven, phased roadmap: grounded in prospective clinical cohort data, progressing from immediate multi-center technical standardization and pragmatic trials, such as minimal residual disease (MRD)-triggered salvage therapy, to mid-term biomarker-driven adjuvant trials and long-term integration into multimodal liquid biopsy ecosystems, aimed at intercepting this execution hub. By reframing CTCs from correlative indicators to actionable therapeutic targets and dynamic sentinels, this framework charts a path toward transforming the management of this recalcitrant systemic disease.

PMID:42150733 | DOI:10.1016/j.phrs.2026.108253

❌