❌

Reading view

Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing

arXiv:2604.02288v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models. While Group Relative Policy Optimization (GRPO) is widely adopted, its coarse credit assignment uniformly penalizes failed rollouts, lacking the token-level focus needed to efficiently address specific deviations. Self-Distillation Policy Optimization (SDPO) addresses this by providing denser, more targeted logit-level supervision that facilitates rapid early improvement, yet it frequently collapses during prolonged training. We trace this late-stage instability to two intrinsic flaws: self-distillation on already-correct samples introduces optimization ambiguity, and the self-teacher's signal reliability progressively degrades. To resolve these issues, we propose Sample-Routed Policy Optimization (SRPO), a unified on-policy framework that routes correct samples to GRPO's reward-aligned reinforcement and failed samples to SDPO's targeted logit-level correction. SRPO further incorporates an entropy-aware dynamic weighting mechanism to suppress high-entropy, unreliable distillation targets while emphasizing confident ones. Evaluated across five benchmarks and two model scales, SRPO achieves both the rapid early improvement of SDPO and the long-horizon stability of GRPO. It consistently surpasses the peak performance of both baselines, raising the five-benchmark average on Qwen3-8B by 3.4% over GRPO and 6.3% over SDPO, while simultaneously yielding moderate response lengths and lowering per-step compute cost by up to 17.2%.
  •  

Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation

arXiv:2603.12793v1 Announce Type: cross Abstract: A recent cutting-edge topic in multimodal modeling is to unify visual comprehension and generation within a single model. However, the two tasks demand mismatched decoding regimes and visual representations, making it non-trivial to jointly optimize within a shared feature space. In this work, we present Cheers, a unified multimodal model that decouples patch-level details from semantic representations, thereby stabilizing semantics for multimodal understanding and improving fidelity for image generation via gated detail residuals. Cheers includes three key components: (i) a unified vision tokenizer that encodes and compresses image latent states into semantic tokens for efficient LLM conditioning, (ii) an LLM-based Transformer that unifies autoregressive decoding for text generation and diffusion decoding for image generation, and (iii) a cascaded flow matching head that decodes visual semantics first and then injects semantically gated detail residuals from the vision tokenizer to refine high-frequency content. Experiments on popular benchmarks demonstrate that Cheers matches or surpasses advanced UMMs in both visual understanding and generation. Cheers also achieves 4x token compression, enabling more efficient high-resolution image encoding and generation. Notably, Cheers outperforms the Tar-1.5B on the popular benchmarks GenEval and MMBench, while requiring only 20% of the training cost, indicating effective and efficient (i.e., 4x token compression) unified multimodal modeling. We will release all code and data for future research.
  •  

Integrated Network Toxicology and Metabolomics Elucidate Mechanisms of Carbosulfan-Induced Respiratory Toxicity in Rats

Int J Mol Sci. 2026 Feb 25;27(5):2170. doi: 10.3390/ijms27052170.

ABSTRACT

Carbosulfan is a widely used carbamate insecticide, yet its mechanisms of respiratory toxicity remain poorly understood. This study integrated network toxicology, untargeted metabolomics, and molecular docking to systematically investigate the potential mechanisms of carbosulfan-induced respiratory toxicity in male Sprague Dawley rats. Rats were administered a single oral dose of carbosulfan (125 or 250 mg/kg) and assessed after 12 h. Exposure resulted in significant pathological lung damage, characterized by disrupted alveolar architecture, inflammatory cell infiltration, and increased serum levels of the pro-inflammatory cytokines IL-6, IL-1Ξ², and TNF-Ξ±. Network toxicology analysis identified 51 potential targets associated with respiratory toxicity, with core targets including SRC, EGFR, PTGS2, CXCL8, CYP3A4, and NR3C1. Enriched pathways were primarily related to neuroactive ligand-receptor interaction, VEGF signaling, and arachidonic acid metabolism. Untargeted metabolomics revealed significant metabolic perturbations in pathways central to antioxidant defense and energy homeostasis, including glutathione metabolism, the tricarboxylic acid cycle, and arginine biosynthesis. Molecular docking confirmed stable in silico binding affinities between carbosulfan and the predicted core targets. Integrative analysis suggests that carbosulfan exposure is associated with respiratory damage, potentially through interconnected mechanisms involving oxidative stress, inflammation, and disruption of cell signaling and metabolic enzyme systems. However, given the acute high-dose nature of the model and the interpretative integration of multi-omics data, these findings should be considered hypothesis-generating. This study provides a novel system-level perspective on carbosulfan-induced respiratory toxicity and highlights key pathways and targets for future validation in chronic exposure models.

PMID:41828400 | PMC:PMC12984169 | DOI:10.3390/ijms27052170

  •  

Integrated Network Toxicology and Metabolomics Elucidate Mechanisms of Carbosulfan-Induced Respiratory Toxicity in Rats

Int J Mol Sci. 2026 Feb 25;27(5):2170. doi: 10.3390/ijms27052170.

ABSTRACT

Carbosulfan is a widely used carbamate insecticide, yet its mechanisms of respiratory toxicity remain poorly understood. This study integrated network toxicology, untargeted metabolomics, and molecular docking to systematically investigate the potential mechanisms of carbosulfan-induced respiratory toxicity in male Sprague Dawley rats. Rats were administered a single oral dose of carbosulfan (125 or 250 mg/kg) and assessed after 12 h. Exposure resulted in significant pathological lung damage, characterized by disrupted alveolar architecture, inflammatory cell infiltration, and increased serum levels of the pro-inflammatory cytokines IL-6, IL-1Ξ², and TNF-Ξ±. Network toxicology analysis identified 51 potential targets associated with respiratory toxicity, with core targets including SRC, EGFR, PTGS2, CXCL8, CYP3A4, and NR3C1. Enriched pathways were primarily related to neuroactive ligand-receptor interaction, VEGF signaling, and arachidonic acid metabolism. Untargeted metabolomics revealed significant metabolic perturbations in pathways central to antioxidant defense and energy homeostasis, including glutathione metabolism, the tricarboxylic acid cycle, and arginine biosynthesis. Molecular docking confirmed stable in silico binding affinities between carbosulfan and the predicted core targets. Integrative analysis suggests that carbosulfan exposure is associated with respiratory damage, potentially through interconnected mechanisms involving oxidative stress, inflammation, and disruption of cell signaling and metabolic enzyme systems. However, given the acute high-dose nature of the model and the interpretative integration of multi-omics data, these findings should be considered hypothesis-generating. This study provides a novel system-level perspective on carbosulfan-induced respiratory toxicity and highlights key pathways and targets for future validation in chronic exposure models.

PMID:41828400 | PMC:PMC12984169 | DOI:10.3390/ijms27052170

  •  

NExT-Guard: Training-Free Streaming Safeguard without Token-Level Labels

arXiv:2603.02219v1 Announce Type: cross Abstract: Large language models are increasingly deployed in streaming scenarios, rendering conventional post-hoc safeguards ineffective as they fail to interdict unsafe content in real-time. While streaming safeguards based on token-level supervised training could address this, they necessitate expensive annotations and suffer from severe overfitting. In this work, we challenge the paradigm that streaming safety must rely on token-level supervised training. Instead, it is an inherent capability of well-trained post-hoc safeguards, as they already encode token-level risk signals in hidden representations. Hence, we introduce NExT-Guard, a training-free framework that achieves streaming safeguards by monitoring interpretable latent features from Sparse Autoencoders (SAEs). It uses pretrained SAEs from publicly available base LLMs, enabling flexible, low-cost deployment without token-level supervision. Experimental results show that NExT-Guard outperforms both post-hoc and streaming safeguards based on supervised training, with superior robustness across models, SAE variants, and risk scenarios. These results make NExT-Guard a universal and scalable paradigm for real-time safety, accelerating the practical deployment of streaming safeguards.
  •  

MultiDiffSense: Diffusion-Based Multi-Modal Visuo-Tactile Image Generation Conditioned on Object Shape and Contact Pose

arXiv:2602.19348v1 Announce Type: cross Abstract: Acquiring aligned visuo-tactile datasets is slow and costly, requiring specialised hardware and large-scale data collection. Synthetic generation is promising, but prior methods are typically single-modality, limiting cross-modal learning. We present MultiDiffSense, a unified diffusion model that synthesises images for multiple vision-based tactile sensors (ViTac, TacTip, ViTacTip) within a single architecture. Our approach uses dual conditioning on CAD-derived, pose-aligned depth maps and structured prompts that encode sensor type and 4-DoF contact pose, enabling controllable, physically consistent multi-modal synthesis. Evaluating on 8 objects (5 seen, 3 novel) and unseen poses, MultiDiffSense outperforms a Pix2Pix cGAN baseline in SSIM by +36.3% (ViTac), +134.6% (ViTacTip), and +64.7% (TacTip). For downstream 3-DoF pose estimation, mixing 50% synthetic with 50% real halves the required real data while maintaining competitive performance. MultiDiffSense alleviates the data-collection bottleneck in tactile sensing and enables scalable, controllable multi-modal dataset generation for robotic applications.
  •  

Eureka-Audio: Triggering Audio Intelligence in Compact Language Models

arXiv:2602.13954v1 Announce Type: cross Abstract: We present Eureka-Audio, a compact yet high-performance audio language model that achieves competitive performance against models that are 4 to 18 times larger across a broad range of audio understanding benchmarks. Despite containing only 1.7B parameters, Eureka-Audio demonstrates strong performance on automatic speech recognition (ASR), audio understanding, and dense audio captioning, matching or surpassing multiple 7B to 30B audio and omni-modal baselines. The model adopts a unified end-to-end architecture composed of a lightweight language backbone, a Whisper-based audio encoder, and a sparsely activated Mixture-of-Experts (MoE) adapter that explicitly accounts for audio heterogeneity and alleviates cross-modal optimization conflicts under limited capacity. To further enhance paralinguistic reasoning, we introduce DataFlux, a closed loop audio instruction data synthesis and verification pipeline that constructs high quality, logically consistent supervision from raw audio. Extensive evaluations across ASR, knowledge reasoning, safety, instruction following, and paralinguistic benchmarks, demonstrate that Eureka-Audio achieves an efficient balance between computational cost and performance. These results establish Eureka Audio as a strong and practical baseline for lightweight audio understanding models.
  •  
❌