❌

Reading view

Discovering Failure Modes in Vision-Language Models using RL

arXiv:2604.04733v1 Announce Type: cross Abstract: Vision-language Models (VLMs), despite achieving strong performance on multimodal benchmarks, often misinterpret straightforward visual concepts that humans identify effortlessly, such as counting, spatial reasoning, and viewpoint understanding. Previous studies manually identified these weaknesses and found that they often stem from deficits in specific skills. However, such manual efforts are costly, unscalable, and subject to human bias, which often overlooks subtle details in favor of salient objects, resulting in an incomplete understanding of a model's vulnerabilities. To address these limitations, we propose a Reinforcement Learning (RL)-based framework to automatically discover the failure modes or blind spots of any candidate VLM on a given data distribution without human intervention. Our framework trains a questioner agent that adaptively generates queries based on the candidate VLM's responses to elicit incorrect answers. Our approach increases question complexity by focusing on fine-grained visual details and distinct skill compositions as training progresses, consequently identifying 36 novel failure modes in which VLMs struggle. We demonstrate the broad applicability of our framework by showcasing its generalizability across various model combinations.
  •  

Dynamic Targetable Extracellular Vesicle Surface Proteins Monitor Depth of Response to CAR T Therapy

Res Sq [Preprint]. 2026 Mar 18:rs.3.rs-8913641. doi: 10.21203/rs.3.rs-8913641/v1.

ABSTRACT

Extracellular vesicles (EVs) represent a promising liquid biopsy platform in multiple myeloma (MM). We developed an MM EV Surface Protein Assay to quantify and dynamically monitor four MM EV subpopulations defined by targetable MM surface proteins (BCMA, CD38, GPRC5D, and CD319) across 336 serial blood samples from 45 relapsed/refractory MM (RRMM) patients treated with anti-BCMA chimeric antigen receptor (CAR) T-cell therapy. All four MM EV subpopulations significantly decreased in 43 patients with initial response, while BCMA+, GPRC5D+, and CD319+ MM EVs increased in 19 patients with progression, and antigen escape was detected by BCMA+ MM EVs. MM EV subpopulations differentiated minimal residual disease (MRD) status and complemented MRD for detecting early relapse before clinical progression. Notably, CD319+ MM EVs were early predictors of progression-free and overall survival in MRD-negative patients. This assay enables noninvasive monitoring of deep response, progression, and antigen escape, and stratifies survival in MRD-negative patients with RRMM.

PMID:41890853 | PMC:PMC13015583 | DOI:10.21203/rs.3.rs-8913641/v1

  •  

Weight Space Representation Learning via Neural Field Adaptation

arXiv:2512.01759v2 Announce Type: replace-cross Abstract: In this work, we investigate the potential of weights to serve as effective representations, focusing on neural fields. Our key insight is that constraining the optimization space through a pre-trained base model and low-rank adaptation (LoRA) can induce structure in weight space. Across reconstruction, generation, and analysis tasks on 2D and 3D data, we find that multiplicative LoRA weights achieve high representation quality while exhibiting distinctiveness and semantic structure. When used with latent diffusion models, multiplicative LoRA weights enable higher-quality generation than existing weight-space methods.
  •  

pFedNavi: Structure-Aware Personalized Federated Vision-Language Navigation for Embodied AI

arXiv:2602.14401v1 Announce Type: cross Abstract: Vision-Language Navigation VLN requires large-scale trajectory instruction data from private indoor environments, raising significant privacy concerns. Federated Learning FL mitigates this by keeping data on-device, but vanilla FL struggles under VLNs' extreme cross-client heterogeneity in environments and instruction styles, making a single global model suboptimal. This paper proposes pFedNavi, a structure-aware and dynamically adaptive personalized federated learning framework tailored for VLN. Our key idea is to personalize where it matters: pFedNavi adaptively identifies client-specific layers via layer-wise mixing coefficients, and performs fine-grained parameter fusion on the selected components (e.g., the encoder-decoder projection and environment-sensitive decoder layers) to balance global knowledge sharing with local specialization. We evaluate pFedNavi on two standard VLN benchmarks, R2R and RxR, using both ResNet and CLIP visual representations. Across all metrics, pFedNavi consistently outperforms the FedAvg-based VLN baseline, achieving up to 7.5% improvement in navigation success rate and up to 7.8% gain in trajectory fidelity, while converging 1.38x faster under non-IID conditions.
  •  
❌