❌

Normal view

SCQ: Stabilizing Conservative Q-Learning with Sigmoid-Bounded Entropy

arXiv:2609.12749v1 Announce Type: new Abstract: Offline-to-online reinforcement learning reduces interaction cost for real-world robot learning but suffers from persistent value estimation instability. Existing methods address this through pessimistic regularization, lower-bound calibration, and architectural normalization, but an overlooked source of instability lies in the entropy formulation: the standard log-entropy term can become negative, destabilizing policy updates. We introduce SCQ (Sigmoid-Bounded Conservative Q-Learning), which replaces this term with a sigmoid-bounded formulation that stays strictly positive. SCQ retains conservative Q regularization and return-based lower-bound calibration, stabilizing policy optimization without sacrificing exploration. We evaluate SCQ on D4RL (Minari) benchmarks under both single-demonstration and standard dataset settings, as well as on simulation and real-world visual tasks. SCQ matches or exceeds baseline performance while exhibiting more stable training dynamics across state-based and visual benchmarks, and transfers to four real-robot platforms including manipulation, wheeled, quadruped, and humanoid systems. A direct clipping intervention that removes negative log-probability contributions, together with gradient-matched positive-score controls, indicates that positivity rather than a particular score shape alone drives much of the improvement. Project website: https://scq-rl.github.io.

Recommendation Retrievers Need Verifiers: Universal Generative Reranking for Sequential Recommendations

arXiv:2609.12270v1 Announce Type: cross Abstract: First-stage recommenders in multi-stage systems produce a ranked candidate list from which a limited prefix is forwarded to downstream rankers. Because each forwarded item must be processed by more expensive ranking stages, this shortlist cannot be arbitrarily large. The first-stage objective is therefore high coverage of relevant items within the forwarded prefix, commonly measured by Recall@$k$. A relevant item may be available deeper in the retrieved list but absent from the shorter prefix that is actually consumed. This paper studies post-hoc verification for promoting such candidates into the consumed shortlist without retraining or replacing the retriever. We introduce a lightweight generative verifier for retrieval models. Given a retriever state and a candidate item, the verifier scores the item through the likelihood of its identifier tokens. It is trained post hoc with next-token cross entropy, requires no sampled negatives or candidate pool during training, and scores only the retriever's top-$K$ candidates at inference. The interface is minimal: the retriever supplies a query state and candidate items, and the item representation can use any fixed tokenization. Across Amazon product recommendation and YaMBDa music recommendation, the same verifier training recipe improves Recall@10 for SASRec, GRU4Rec, NextItNet, and MiniOneRec. Ablations show that the improvements are not explained solely by injecting item-content features into the retriever, supporting verification as a post-hoc output-side adaptation mechanism.

High-salt diet in macrophage-associated metabolic disorders: Mechanisms and therapeutic implications

Chin Med J (Engl). 2026 May 19. doi: 10.1097/CM9.0000000000004098. Online ahead of print.

ABSTRACT

High-salt diet (HSD) has emerged as a prevalent environmental factor that exacerbates chronic inflammation and insulin resistance in obesity-associated type 2 diabetes (T2D) by modulating macrophage polarization, metabolic reprogramming, and epigenetic imprinting. Current evidence demonstrates that HSD activates p38/mitogen-activated protein kinase (MAPK), nuclear factor kappa-B (NF-κB), and NOD-like receptor family pyrin domain containing 3 (NLRP3) inflammasome signaling pathways, by which it drives macrophage polarization toward a proinflammatory M1 phenotype while inducing a glycolysis-dominant metabolic shift, thereby establishing a persistent "metabolic memory". Moreover, HSD orchestrates metabolic memory in macrophages through coordinated epigenetic machinery, including histone modifications (Trimethylation of histone H3 at lysine 4 [H3K4me3] and Acetylation of histone H3 at lysine 27 [H3K27ac]), DNA methylation, and noncoding RNAs (e.g., long non-coding RNA MALAT1 and miR-155), leading to sustained inflammatory phenotypes. In multiple metabolic organs (e.g., adipose tissue, liver, pancreas, and gut), the HSD-macrophage axis aggravates systemic insulin resistance through shared proinflammatory signaling and other tissue-specific mechanisms. Most importantly, therapeutic strategies targeting the NLRP3 inflammasome, metabolic pathways, and epigenetic alterations offer novel approaches for managing metabolic inflammation. Future investigations are encouraged to leverage lineage tracing, single-cell sequencing, and spatial multi-omics technologies to advance the development of precision medicine for macrophage-associated metabolic disorders.

PMID:42156155 | DOI:10.1097/CM9.0000000000004098

WIMLE: Uncertainty-Aware World Models with IMLE for Sample-Efficient Continuous Control

arXiv:2602.14351v2 Announce Type: replace-cross Abstract: Model-based reinforcement learning promises strong sample efficiency but often underperforms in practice due to compounding model error, unimodal world models that average over multi-modal dynamics, and overconfident predictions that bias learning. We introduce WIMLE, a model-based method that extends Implicit Maximum Likelihood Estimation (IMLE) to the model-based RL framework to learn stochastic, multi-modal world models without iterative sampling and to estimate predictive uncertainty via ensembles and latent sampling. During training, WIMLE weights each synthetic transition by its predicted confidence, preserving useful model rollouts while attenuating bias from uncertain predictions and enabling stable learning. Across $40$ continuous-control tasks spanning DeepMind Control, MyoSuite, and HumanoidBench, WIMLE achieves superior sample efficiency and competitive or better asymptotic performance than strong model-free and model-based baselines. Notably, on the challenging Humanoid-run task, WIMLE improves sample efficiency by over $50$\% relative to the strongest competitor, and on HumanoidBench it solves $8$ of $14$ tasks (versus $4$ for BRO and $5$ for SimbaV2). These results highlight the value of IMLE-based multi-modality and uncertainty-aware weighting for stable model-based RL.

MindCube: Spatial Mental Modeling from Limited Views

arXiv:2506.21458v2 Announce Type: replace Abstract: Can Vision-Language Models (VLMs) imagine the full scene from just a few views, like humans do? Humans form spatial mental models naturally, internal representations of unseen space, to reason about layout, perspective, and motion. Our MindCube benchmark with 21,154 questions across 3,268 images exposes this critical gap, where existing VLMs exhibit near-random performance. Using MindCube, we systematically evaluate how well VLMs build robust spatial mental models through representing positions (cognitive mapping), orientations (perspective-taking), and dynamics (mental simulation for "what-if" movements). We then explore three approaches to help approximate spatial mental models in VLMs, focusing on incorporating unseen intermediate views, natural language reasoning chains, and cognitive maps. The significant improvement comes from a synergistic approach, "map-then-reason", that jointly trains the model to first generate a cognitive map and then reason upon it. By training models to reason over these internal maps, we boosted accuracy from 37.8% to 57.8% (+20.0%). Adding reinforcement learning pushed performance even further to 61.3% (+23.5%). Our key insight is that such scaffolding of spatial mental models, actively constructing and utilizing internal structured spatial representations with flexible reasoning processes, significantly improves understanding of unobservable space.

Robust transcriptomic hallmarks targeting intratumor heterogeneity in intrahepatic cholangiocarcinoma

Cell Rep Med. 2026 Mar 30:102708. doi: 10.1016/j.xcrm.2026.102708. Online ahead of print.

ABSTRACT

Intratumor heterogeneity (ITH) undermines transcriptome-based stratification in intrahepatic cholangiocarcinoma (iCCA). Here, we integrate multi-omics data from multi-region, single-region, and single-cell RNA sequencing cohorts to systematically characterize gene expression ITH. We uncover that immune and stromal heterogeneity are primary drivers of ITH, leading to misclassification of a median 27.8% of tumors by existing subtyping systems. To overcome this, we identify a low-intratumor-heterogeneity/high-intertumor-variability (LIHV) gene set and develop an ITH-insensitive classification system defining five subgroups: inflammatory (SI), metabolic (SII), atypical (SIII-1), immune-silent (SIII-2), and neurodegenerative (SIII-3). These subgroups exhibit distinct clinical outcomes, molecular features, immune landscapes, and therapeutic vulnerabilities. GPRC5A and VTCN1 serve as robust immunohistochemical biomarkers for SI and SIII tumors, while serum CEA and CA19-9 identify inflammatory iCCA. Therapeutically, HSP90 inhibition synergizes with anti-PD1 in inflammatory iCCA, whereas combined anti-PD1 and anti-TIM3 suppresses neurodegenerative iCCA. Collectively, our study provides a robust molecular framework and actionable therapeutic strategies for iCCA.

PMID:41916296 | DOI:10.1016/j.xcrm.2026.102708

Single-cell multiomics uncovers an endothelial mechanosensitive PIEZO1-IL-33 axis driving pulmonary fibrosis

Nat Commun. 2026 Mar 20;17(1):2655. doi: 10.1038/s41467-026-70193-w.

ABSTRACT

Pulmonary fibrosis represents a progressive interstitial lung disease marked by excessive extracellular matrix deposition and architectural distortion. Vascular endothelial cells critically contribute to fibrogenesis through paracrine secretion of pro-fibrotic mediators, yet their mechanobiological regulation remains elusive. Using integrated single-cell multi-omics profiling of human pulmonary fibrosis specimens and experimental fibrosis models induced by bleomycin or silica, we identify mechanosensitive Piezo1 upregulation in Endothelial cells as a hallmark of fibrotic progression. Endothelial-specific Piezo1 knockout significantly attenuates Bleomycin-induced fibrotic remodeling in male mice, establishing its pathogenic necessity. Mechanistically, PIEZO1 activation promotes pulmonary fibrosis development via CAPN2-mediated STAT3 phosphorylation, which may regulate the secretion of the pro-fibrotic molecule interleukin-33. These findings suggest that the endothelial PIEZO1-CAPN2-STAT3-IL33 axis is a potential therapeutic target for PF intervention.

PMID:41862476 | PMC:PMC13004862 | DOI:10.1038/s41467-026-70193-w

Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion

arXiv:2603.03485v1 Announce Type: cross Abstract: Recent video diffusion models have achieved impressive capabilities as large-scale generative world models. However, these models often struggle with fine-grained physical consistency, exhibiting physically implausible dynamics over time. In this work, we present \textbf{Phys4D}, a pipeline for learning physics-consistent 4D world representations from video diffusion models. Phys4D adopts \textbf{a three-stage training paradigm} that progressively lifts appearance-driven video diffusion models into physics-consistent 4D world representations. We first bootstrap robust geometry and motion representations through large-scale pseudo-supervised pretraining, establishing a foundation for 4D scene modeling. We then perform physics-grounded supervised fine-tuning using simulation-generated data, enforcing temporally consistent 4D dynamics. Finally, we apply simulation-grounded reinforcement learning to correct residual physical violations that are difficult to capture through explicit supervision. To evaluate fine-grained physical consistency beyond appearance-based metrics, we introduce a set of \textbf{4D world consistency evaluation} that probe geometric coherence, motion stability, and long-horizon physical plausibility. Experimental results demonstrate that Phys4D substantially improves fine-grained spatiotemporal and physical consistency compared to appearance-driven baselines, while maintaining strong generative performance. Our project page is available at https://sensational-brioche-7657e7.netlify.app/

WIMLE: Uncertainty-Aware World Models with IMLE for Sample-Efficient Continuous Control

arXiv:2602.14351v1 Announce Type: cross Abstract: Model-based reinforcement learning promises strong sample efficiency but often underperforms in practice due to compounding model error, unimodal world models that average over multi-modal dynamics, and overconfident predictions that bias learning. We introduce WIMLE, a model-based method that extends Implicit Maximum Likelihood Estimation (IMLE) to the model-based RL framework to learn stochastic, multi-modal world models without iterative sampling and to estimate predictive uncertainty via ensembles and latent sampling. During training, WIMLE weights each synthetic transition by its predicted confidence, preserving useful model rollouts while attenuating bias from uncertain predictions and enabling stable learning. Across $40$ continuous-control tasks spanning DeepMind Control, MyoSuite, and HumanoidBench, WIMLE achieves superior sample efficiency and competitive or better asymptotic performance than strong model-free and model-based baselines. Notably, on the challenging Humanoid-run task, WIMLE improves sample efficiency by over $50$\% relative to the strongest competitor, and on HumanoidBench it solves $8$ of $14$ tasks (versus $4$ for BRO and $5$ for SimbaV2). These results highlight the value of IMLE-based multi-modality and uncertainty-aware weighting for stable model-based RL.
❌