❌

Normal view

RAE-AR: Taming Autoregressive Models with Representation Autoencoders

arXiv:2604.01545v1 Announce Type: new Abstract: The latent space of generative modeling is long dominated by the VAE encoder. The latents from the pretrained representation encoders (e.g., DINO, SigLIP, MAE) are previously considered inappropriate for generative modeling. Recently, RAE method lights the hope and reveals that the representation autoencoder can also achieve competitive performance as the VAE encoder. However, the integration of representation autoencoder into continuous autoregressive (AR) models, remains largely unexplored. In this work, we investigate the challenges of employing high-dimensional representation autoencoders within the AR paradigm, denoted as \textit{RAE-AR}. We focus on the unique properties of AR models and identify two primary hurdles: complex token-wise distribution modeling and the high-dimensionality amplified training-inference gap (exposure bias). To address these, we introduce token simplification via distribution normalization to ease modeling difficulty and improve convergence. Furthermore, we enhance prediction robustness by incorporating Gaussian noise injection during training to mitigate exposure bias. Our empirical results demonstrate that these modifications substantially bridge the performance gap, enabling representation autoencoder to achieve results comparable to traditional VAEs on AR models. This work paves the way for a more unified architecture across visual understanding and generative modeling.

Leptomeningeal metastatic cancer cells induce a permissive choroid plexus vasculature through extracellular-vesicle-derived 5-HIAA signaling

Nature Cancer, Published online: 03 April 2026; doi:10.1038/s43018-026-01145-y

Huang, Hou, Yang et al. demonstrate that leptomeningeal metastatic cells favor the formation of a premetastatic niche by remodeling the choroid plexus vasculature through the serotonin metabolite 5-hydroxyindoleacetic acid, which signals into endothelial cells through the aryl hydrocarbon receptor.

Low-dose intestinal irradiation enhances the efficacy and prognosis of PD-1 blockade in metastatic non-small cell lung cancer

Clin Cancer Res. 2026 Mar 18. doi: 10.1158/1078-0432.CCR-25-4153. Online ahead of print.

ABSTRACT

PURPOSE: Intestinal low-dose irradiation (ILDR) may enhance immunotherapy efficacy by modulating the gut microbiota and metabolism; however, its role in metastatic non-small cell lung cancer (mNSCLC), particularly in the first-line setting, remains unclear.

EXPERIMENTAL DESIGN: This multicenter retrospective and prospective study included mNSCLC patients receiving first- and second-line programmed cell death protein 1 (PD-1) inhibitors along with abdominopelvic radiotherapy between 2018 and 2025. Patients were stratified by the mean intestinal radiation dose into <1 Gy, 1-3 Gy, and >3 Gy groups and treatment outcomes were compared. The blood and fecal samples were subjected to multi-omics profiling.

RESULTS: g>309 patients were included in the retrospective analysis. Optimal efficacy was observed with a small intestinal mean radiation dose (SIMRD) of 1-3 Gy, showing longer progression-free survival (PFS, 10.2 months) and overall survival (OS, 22.8 months) (P < 0.01), which was consistent across subgroups. Compared with 1-3 Gy, SIMRD >3 Gy (Hazard ratio [HR] = 4.87, P < 0.001) and <1 Gy (HR = 1.85, P < 0.001) independently predicted worse OS. Prospective results confirmed the best disease control rate (P = 0.041) and PFS (P = 0.046) with SIMRD of 1-3 Gy. Responders were enriched in Bacillota, Clostridia, and indole derivatives, particularly indole-3-carboxylic acid. Moreover, the 1-3 Gy group exhibited increased circulating macrophage inflammatory protein-3α and reduced circulating α4β7+ regulatory T cells.

CONCLUSIONS: ILDR influences the efficacy of PD-1 blockade in patients with mNSCLC, particularly when SIMRD is maintained within the 1-3 Gy range, likely through modulation of the gut microbiota-metabolite-immune axis.

PMID:41849236 | DOI:10.1158/1078-0432.CCR-25-4153

CAPT: Confusion-Aware Prompt Tuning for Reducing Vision-Language Misalignment

arXiv:2603.02557v1 Announce Type: cross Abstract: Vision-language models like CLIP have achieved remarkable progress in cross-modal representation learning, yet suffer from systematic misclassifications among visually and semantically similar categories. We observe that such confusion patterns are not random but persistently occur between specific category pairs, revealing the model's intrinsic bias and limited fine-grained discriminative ability. To address this, we propose CAPT, a Confusion-Aware Prompt Tuning framework that enables models to learn from their own misalignment. Specifically, we construct a Confusion Bank to explicitly model stable confusion relationships across categories and misclassified samples. On this basis, we introduce a Semantic Confusion Miner (SEM) to capture global inter-class confusion through semantic difference and commonality prompts, and a Sample Confusion Miner (SAM) to retrieve representative misclassified instances from the bank and capture sample-level cues through a Diff-Manner Adapter that integrates global and local contexts. To further unify confusion information across different granularities, a Multi-Granularity Difference Expert (MGDE) module is designed to jointly leverage semantic- and sample-level experts for more robust confusion-aware reasoning. Extensive experiments on 11 benchmark datasets demonstrate that our method significantly reduces confusion-induced errors while enhancing the discriminability and generalization of both base and novel classes, successfully resolving 50.72 percent of confusable sample pairs. Code will be released at https://github.com/greatest-gourmet/CAPT.
❌