❌

Reading view

Countering Catastrophic Forgetting of Large Language Models for Better Instruction Following via Weight-Space Model Merging

arXiv:2604.01538v1 Announce Type: cross Abstract: Large language models have been adopted in the medical domain for clinical documentation to reduce clinician burden. However, studies have reported that LLMs often "forget" a significant amount of instruction-following ability when fine-tuned using a task-specific medical dataset, a critical challenge in adopting general-purpose LLMs for clinical applications. This study presents a model merging framework to efficiently adapt general-purpose LLMs to the medical domain by countering this forgetting issue. By merging a clinical foundation model (GatorTronLlama) with a general instruct model (Llama-3.1-8B-Instruct) via interpolation-based merge methods, we seek to derive a domain-adapted model with strong performance on clinical tasks while retaining instruction-following ability. Comprehensive evaluation across medical benchmarks and five clinical generation tasks (e.g., radiology and discharge summarization) shows that merged models can effectively mitigate catastrophic forgetting, preserve clinical domain expertise, and retain instruction-following ability. In addition, our model merging strategies demonstrate training efficiency, achieving performance on par with fully fine-tuned baselines under severely constrained supervision (e.g., 64-shot vs. 256-shot). Consequently, weight-space merging constitutes a highly scalable solution for adapting open-source LLMs to clinical applications, facilitating broader deployment in resource-constrained healthcare environments.
  •  

Catgut implantation at acupoints improves anti-PD-1 inhibitor efficacy in lung cancer by inducing immune responses and remodeling the tumor microenvironment

Cancer Immunol Immunother. 2026 Mar 31;75(4):126. doi: 10.1007/s00262-026-04368-1.

ABSTRACT

While anti-programmed death-1 (anti-PD-1) therapy has revolutionized lung cancer treatment, its efficacy remains limited by an immunosuppressive tumor microenvironment (TME). We therefore investigated whether combining anti-PD-1 inhibitor with catgut embedding at the Zusanli acupoint (CIAA) could enhance anti-tumor immunity by reprogramming the TME in a lung cancer mouse model. Combining in vivo tumor monitoring, multi-parametric immune profiling (flow cytometry, IHC, ELISA), and multi-omics analyses (transcriptomics and metabolomics), we found that the combination therapy was associated with enhanced tumor growth inhibition. This effect correlated with a comprehensive TME transformation: conversion to an immunologically active state with increased effector immune cell infiltration (CD8⁺ T, CD4⁺ T, B cells, macrophages) and decreased regulatory T cells, coupled with suppression of pro-tumorigenic factors (VEGF, IL-6). Integrated omics analysis suggests that the combined treatment may modulate tumor-stroma interaction pathways (e.g., PI3K-Akt, focal adhesion) and rewire immunometabolic networks (e.g., tryptophan metabolism). Our study provides hypothesis-generating correlative data positioning CIAA as a potential adjunct capable of remodeling the TME to potentiate anti-PD-1 therapy in lung cancer.

PMID:41915222 | PMC:PMC13038699 | DOI:10.1007/s00262-026-04368-1

  •  

Catgut implantation at acupoints improves anti-PD-1 inhibitor efficacy in lung cancer by inducing immune responses and remodeling the tumor microenvironment

Cancer Immunol Immunother. 2026 Mar 31;75(4):126. doi: 10.1007/s00262-026-04368-1.

ABSTRACT

While anti-programmed death-1 (anti-PD-1) therapy has revolutionized lung cancer treatment, its efficacy remains limited by an immunosuppressive tumor microenvironment (TME). We therefore investigated whether combining anti-PD-1 inhibitor with catgut embedding at the Zusanli acupoint (CIAA) could enhance anti-tumor immunity by reprogramming the TME in a lung cancer mouse model. Combining in vivo tumor monitoring, multi-parametric immune profiling (flow cytometry, IHC, ELISA), and multi-omics analyses (transcriptomics and metabolomics), we found that the combination therapy was associated with enhanced tumor growth inhibition. This effect correlated with a comprehensive TME transformation: conversion to an immunologically active state with increased effector immune cell infiltration (CD8⁺ T, CD4⁺ T, B cells, macrophages) and decreased regulatory T cells, coupled with suppression of pro-tumorigenic factors (VEGF, IL-6). Integrated omics analysis suggests that the combined treatment may modulate tumor-stroma interaction pathways (e.g., PI3K-Akt, focal adhesion) and rewire immunometabolic networks (e.g., tryptophan metabolism). Our study provides hypothesis-generating correlative data positioning CIAA as a potential adjunct capable of remodeling the TME to potentiate anti-PD-1 therapy in lung cancer.

PMID:41915222 | DOI:10.1007/s00262-026-04368-1

  •  

A monocyte-centered framework for predicting immunochemotherapy efficacy in lung squamous cell carcinoma patients

EMBO Mol Med. 2026 Mar 30. doi: 10.1038/s44321-026-00410-y. Online ahead of print.

ABSTRACT

Lung cancer is the leading cause of cancer-related mortality worldwide, with lung squamous cell carcinoma (LUSC) comprising 20-30% of cases. Immunochemotherapy (IC) is the standard first-line treatment for advanced LUSC, yet reliable predictors of therapeutic response remain unavailable. Using single-cell multi-omics profiling of paired pre- and post-treatment tumor and blood samples, we observed that patients responding to IC exhibited significantly higher baseline levels of peripheral blood monocytes, tumor-infiltrating classical monocytes, and APOBEC3A+ monocytes across both compartments compared with non-responders. These associations were independently validated in additional cohorts using routine complete blood count testing and multiplex immunofluorescence analysis of native tumor tissues. Our findings reveal monocyte-related parameters as clinically accessible indicators that link systemic immunity with the tumor microenvironment and hold promise for predicting IC responsiveness in patients with LUSC.

PMID:41912871 | DOI:10.1038/s44321-026-00410-y

  •  

A monocyte-centered framework for predicting immunochemotherapy efficacy in lung squamous cell carcinoma patients

EMBO Mol Med. 2026 Mar 30. doi: 10.1038/s44321-026-00410-y. Online ahead of print.

ABSTRACT

Lung cancer is the leading cause of cancer-related mortality worldwide, with lung squamous cell carcinoma (LUSC) comprising 20-30% of cases. Immunochemotherapy (IC) is the standard first-line treatment for advanced LUSC, yet reliable predictors of therapeutic response remain unavailable. Using single-cell multi-omics profiling of paired pre- and post-treatment tumor and blood samples, we observed that patients responding to IC exhibited significantly higher baseline levels of peripheral blood monocytes, tumor-infiltrating classical monocytes, and APOBEC3A+ monocytes across both compartments compared with non-responders. These associations were independently validated in additional cohorts using routine complete blood count testing and multiplex immunofluorescence analysis of native tumor tissues. Our findings reveal monocyte-related parameters as clinically accessible indicators that link systemic immunity with the tumor microenvironment and hold promise for predicting IC responsiveness in patients with LUSC.

PMID:41912871 | DOI:10.1038/s44321-026-00410-y

  •  

Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance

arXiv:2602.01047v3 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) can reason from image-text inputs and perform well in various multimodal tasks. Despite this success, they are affected by language priors and often produce hallucinations. Hallucinations denote generated content that is grammatically and syntactically coherent, yet bears no match or direct relevance to visual input. To address this problem, we propose Residual Decoding (ResDec). It is a novel training-free method that uses historical information to aid decoding. The method relies on the internal implicit reasoning mechanism and token logits evolution mechanism of LVLMs to correct biases. Extensive experiments demonstrate that ResDec effectively suppresses hallucinations induced by language priors, significantly improves visual grounding, and reduces object hallucinations. In addition to mitigating hallucinations, ResDec also performs exceptionally well on comprehensive LVLM benchmarks, highlighting its broad applicability.
  •  

When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning

arXiv:2603.21289v2 Announce Type: replace-cross Abstract: Recent progress in multimodal large language models has led to strong performance on reasoning tasks, but these improvements largely rely on high-quality annotated data or teacher-model distillation, both of which are costly and difficult to scale. To address this, we propose an unsupervised self-evolution training framework for multimodal reasoning that achieves stable performance improvements without using human-annotated answers or external reward models. For each input, we sample multiple reasoning trajectories and jointly model their within group structure. We use the Actor's self-consistency signal as a training prior, and introduce a bounded Judge based modulation to continuously reweight trajectories of different quality. We further model the modulated scores as a group level distribution and convert absolute scores into relative advantages within each group, enabling more robust policy updates. Trained with Group Relative Policy Optimization (GRPO) on unlabeled data, our method consistently improves reasoning performance and generalization on five mathematical reasoning benchmarks, offering a scalable path toward self-evolving multimodal models. The code are available at https://github.com/OPPO-Mente-Lab/LLM-Self-Judge.
  •  

Effects of multisensory stimulation based on immersive virtual reality in postoperative neuropsychiatric recovery after gynecological laparoscopy

npj Digital Medicine, Published online: 24 March 2026; doi:10.1038/s41746-026-02515-7

Effects of multisensory stimulation based on immersive virtual reality in postoperative neuropsychiatric recovery after gynecological laparoscopy
  •  

Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views

arXiv:2510.18632v4 Announce Type: replace-cross Abstract: Though recent advances in vision-language models (VLMs) have achieved remarkable progress across a wide range of multimodal tasks, understanding 3D spatial relationships from limited views remains a significant challenge. Previous reasoning methods typically rely on pure text (e.g., topological cognitive maps) or on 2D visual cues. However, their limited representational capacity hinders performance in specific tasks that require 3D spatial imagination. To address this limitation, we propose 3DThinker, a framework that can effectively exploits the rich geometric information embedded within images while reasoning, like humans do. Our framework is the first to enable 3D mentaling during reasoning without any 3D prior input, and it does not rely on explicitly labeled 3D data for training. Specifically, our training consists of two stages. First, we perform supervised training to align the 3D latent generated by VLM while reasoning with that of a 3D foundation model (e.g., VGGT). Then, we optimize the entire reasoning trajectory solely based on outcome signals, thereby refining the underlying 3D mentaling. Extensive experiments across multiple benchmarks show that 3DThinker consistently outperforms strong baselines and offers a new perspective toward unifying 3D representations into multimodal reasoning. Our code is available at https://github.com/zhangquanchen/3DThinker.
  •  

Multimodal Continual Learning with MLLMs from Multi-scenario Perspectives

arXiv:2511.18507v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) deployed on devices must adapt to continuously changing visual scenarios such as variations in background and perspective, to effectively perform complex visual tasks. To investigate catastrophic forgetting under real-world scenario shifts, we construct a multimodal visual understanding dataset (MSVQA), covering four distinct scenarios and perspectives: high-altitude, underwater, low-altitude, and indoor environments. Furthermore, we propose UNIFIER (mUltimodal coNtInual learning with MLLMs From multi-scenarIo pERspectives), a continual learning (CL) framework designed to address visual discrepancies while learning different scenarios. Compared to existing CL methods, UNIFIER enables knowledge accumulation within the same scenario and mutual enhancement across different scenarios via Vision Representation Expansion (VRE) and Vision Consistency Constraint (VCC). Experimental results show that UNIFIER improves the last-step VQA scores by 2.70%~10.62% and the last-step F1 scores by 3.40%~7.69% compared to the state-of-the-art method, QUAD, in 20-step cross-scenario continual learning tasks. MSVQA dataset is available at https://huggingface.co/datasets/Kaij00/MSVQA.
  •  

Risk-adaptive therapy guided by dynamic ctDNA in nasopharyngeal carcinoma

Nature, Published online: 11 March 2026; doi:10.1038/s41586-026-10244-w

A clinical trial testing whether monitoring ctDNA clearance during treatment for nasopharyngeal cancer could be used to inform decisions about an individual’s subsequent therapeutic programme shows promising results.
  •  

Attn-QAT: 4-Bit Attention With Quantization-Aware Training

arXiv:2603.00040v2 Announce Type: replace-cross Abstract: Achieving reliable 4-bit attention is a prerequisite for end-to-end FP4 computation on emerging FP4-capable GPUs, yet attention remains the main obstacle due to FP4's tiny dynamic range and attention's heavy-tailed activations. This paper presents the first systematic study of 4-bit quantization-aware training (QAT) for attention. We find that "drop-in" QAT, which naively combines an FP4 forward pass with a high-precision Flash Attention (FA)-style backward pass, leads to training instability. We identify two key principles for stable FP4 attention: (1) matching low-precision recomputation of attention scores in the backward pass, and (2) resolving implicit precision assumptions in FA's gradient calculation. Based on these insights, we propose Attn-QAT and implement fused Triton kernels for training as well as FP4 inference kernels. Across diffusion and language models, Attn-QAT recovers the quality drop from FP4 attention without explicit outlier-mitigation heuristics used in prior FP4 attention, and delivers up to a 1.5x speedup on an RTX 5090. Video demos can be found at https://drive.google.com/drive/folders/190F6xbBDUF2kGQYIcXBt3ehSYij5jlim?usp=sharing.
  •  

ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training

arXiv:2603.04385v2 Announce Type: replace-cross Abstract: Feed-forward transformer models have driven rapid progress in 3D vision, but state-of-the-art methods such as VGGT and $\pi^3$ have a computational cost that scales quadratically with the number of input images, making them inefficient when applied to large image collections. Sequential-reconstruction approaches reduce this cost but sacrifice reconstruction quality. We introduce ZipMap, a stateful feed-forward model that achieves linear-time, bidirectional 3D reconstruction while matching or surpassing the accuracy of quadratic-time methods. ZipMap employs test-time training layers to zip an entire image collection into a compact hidden scene state in a single forward pass, enabling reconstruction of over 700 frames in under 10 seconds on a single H100 GPU, more than $20\times$ faster than state-of-the-art methods such as VGGT. Moreover, we demonstrate the benefits of having a stateful representation in real-time scene-state querying and its extension to sequential streaming reconstruction.
  •  

<i>KRAS</i>-extrachromosomal DNA drives intratumoral heterogeneity in gastric cancer

Oncogene, Published online: 05 March 2026; doi:10.1038/s41388-026-03713-z

KRAS-extrachromosomal DNA drives intratumoral heterogeneity in gastric cancer
  •  

Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

arXiv:2603.04128v1 Announce Type: cross Abstract: Developing Audio-Visual Large Language Models (AV-LLMs) for unified scene understanding is pivotal in multimodal intelligence. While instruction tuning enables pre-trained models with multi-task abilities, we observe that conventional multi-task unification methods often suffer from severe negative transfer, where nearly 55% of tasks degrade compared to single-task training. We attribute this phenomenon to audio-visual task heterogeneity, characterized by disparate task granularity and divergent capability demands, which lead to negative interference under joint training. To tackle this, we present Crab$^{+}$, a scalable and unified audio-visual scene understanding model that addresses task heterogeneity through explicit cooperation from both data and model perspectives. On the data side, we introduce AV-UIE v2, a comprehensive Audio-Visual Unified Instruction-tuning dataset with Explicit reasoning processes. It contains approximately 222K samples spanning 17 datasets and 7 tasks, enabling the model to capture cross-task relationships at different levels of granularity. On the model side, we design a unified interface to align heterogeneous task formulations, and propose Interaction-aware LoRA (I-LoRA), which explicitly models inter-task relationships via dynamic routing to coordinate distinct audio-visual interaction patterns, mitigating parameter interference. Extensive experiments show Crab$^{+}$ covers broader tasks than existing unified models while outperforming specialized models on various benchmarks. We successfully reverse the negative transfer trend, achieving positive transfer where multi-task learning surpasses single-task baselines in nearly 88% of tasks. These results hold across diverse AV-LLM paradigms and are validated through in-depth visualization, positioning Crab$^{+}$ as a robust step towards holistic audio-visual scene understanding.
  •  

ZipMap: Linear-Time Stateful 3D Reconstruction with Test-Time Training

arXiv:2603.04385v1 Announce Type: cross Abstract: Feed-forward transformer models have driven rapid progress in 3D vision, but state-of-the-art methods such as VGGT and $\pi^3$ have a computational cost that scales quadratically with the number of input images, making them inefficient when applied to large image collections. Sequential-reconstruction approaches reduce this cost but sacrifice reconstruction quality. We introduce ZipMap, a stateful feed-forward model that achieves linear-time, bidirectional 3D reconstruction while matching or surpassing the accuracy of quadratic-time methods. ZipMap employs test-time training layers to zip an entire image collection into a compact hidden scene state in a single forward pass, enabling reconstruction of over 700 frames in under 10 seconds on a single H100 GPU, more than $20\times$ faster than state-of-the-art methods such as VGGT. Moreover, we demonstrate the benefits of having a stateful representation in real-time scene-state querying and its extension to sequential streaming reconstruction.
  •  

RubricBench: Aligning Model-Generated Rubrics with Human Standards

arXiv:2603.01562v2 Announce Type: replace Abstract: As Large Language Model (LLM) alignment evolves from simple completions to complex, highly sophisticated generation, Reward Models are increasingly shifting toward rubric-guided evaluation to mitigate surface-level biases. However, the community lacks a unified benchmark to assess this evaluation paradigm, as existing benchmarks lack both the discriminative complexity and the ground-truth rubric annotations required for rigorous analysis. To bridge this gap, we introduce RubricBench, a curated benchmark with 1,147 pairwise comparisons specifically designed to assess the reliability of rubric-based evaluation. Our construction employs a multi-dimensional filtration pipeline to target hard samples featuring nuanced input complexity and misleading surface bias, augmenting each with expert-annotated, atomic rubrics derived strictly from instructions. Comprehensive experiments reveal a substantial capability gap between human-annotated and model-generated rubrics, indicating that even state-of-the-art models struggle to autonomously specify valid evaluation criteria, lagging considerably behind human-guided performance.
  •  

NeuroWise: A Multi-Agent LLM "Glass-Box" System for Practicing Double-Empathy Communication with Autistic Partners

arXiv:2602.18962v2 Announce Type: replace-cross Abstract: The double empathy problem frames communication difficulties between neurodivergent and neurotypical individuals as arising from mutual misunderstanding, yet most interventions focus on autistic individuals. We present NeuroWise, a multi-agent LLM-based coaching system that supports neurotypical users through stress visualization, interpretation of internal experiences, and contextual guidance. In a between-subjects study (N=30), NeuroWise was rated as helpful by all participants and showed a significant condition-time effect on deficit-based attributions (p=0.02): NeuroWise users reduced deficit framing, while baseline users shifted toward blaming autistic "deficits" after difficult interactions. NeuroWise users also completed conversations more efficiently (37% fewer turns, p=0.03). These findings suggest that AI-based interpretation can support attributional change by helping users recognize communication challenges as mutual.
  •  

Give Users the Wheel: Towards Promptable Recommendation Paradigm

arXiv:2602.18929v1 Announce Type: cross Abstract: Conventional sequential recommendation models have achieved remarkable success in mining implicit behavioral patterns. However, these architectures remain structurally blind to explicit user intent: they struggle to adapt when a user's immediate goal (e.g., expressed via a natural language prompt) deviates from their historical habits. While Large Language Models (LLMs) offer the semantic reasoning to interpret such intent, existing integration paradigms force a dilemma: LLM-as-a-recommender paradigm sacrifices the efficiency and collaborative precision of ID-based retrieval, while Reranking methods are inherently bottlenecked by the recall capabilities of the underlying model. In this paper, we propose Decoupled Promptable Sequential Recommendation (DPR), a model-agnostic framework that empowers conventional sequential backbones to natively support Promptable Recommendation, the ability to dynamically steer the retrieval process using natural language without abandoning collaborative signals. DPR modulates the latent user representation directly within the retrieval space. To achieve this, we introduce a Fusion module to align the collaborative and semantic signals, a Mixture-of-Experts (MoE) architecture that disentangles the conflicting gradients from positive and negative steering, and a three-stage training strategy that progressively aligns the semantic space of prompts with the collaborative space. Extensive experiments on real-world datasets demonstrate that DPR significantly outperforms state-of-the-art baselines in prompt-guided tasks while maintaining competitive performance in standard sequential recommendation scenarios.
  •  

NeuroWise: A Multi-Agent LLM "Glass-Box" System for Practicing Double-Empathy Communication with Autistic Partners

arXiv:2602.18962v1 Announce Type: cross Abstract: The double empathy problem frames communication difficulties between neurodivergent and neurotypical individuals as arising from mutual misunderstanding, yet most interventions focus on autistic individuals. We present NeuroWise, a multi-agent LLM-based coaching system that supports neurotypical users through stress visualization, interpretation of internal experiences, and contextual guidance. In a between-subjects study (N=30), NeuroWise was rated as helpful by all participants and showed a significant condition-time effect on deficit-based attributions (p=0.02): NeuroWise users reduced deficit framing, while baseline users shifted toward blaming autistic "deficits" after difficult interactions. NeuroWise users also completed conversations more efficiently (37% fewer turns, p=0.03). These findings suggest that AI-based interpretation can support attributional change by helping users recognize communication challenges as mutual.
  •  
❌