❌

Normal view

Viral gene replication enhances AAV vector quality and reduces manufacturing costs

Liu and colleagues developed a robust in cellulo plasmid DNA replication system in human cells for replicating plasmid-borne adeno-associated virus (AAV) Rep/Cap genes during recombinant AAV (rAAV) production. This new approach not only enables a 10- to 20-fold plasmid reduction to significantly lower manufacturing costs but also substantially enhances rAAV potency, titer, and purity.

Multi-omics integration identifies ribosome biogenesis-active macrophage subpopulation and its key gene GNL2 in driving liver hepatocellular carcinoma progression and mechanisms

Cancer Cell Int. 2026 May 14. doi: 10.1186/s12935-026-04330-2. Online ahead of print.

ABSTRACT

BACKGROUND: Liver hepatocellular carcinoma (LIHC) is a common malignancy, yet the core genes driving its progression and potential therapeutic targets remain insufficiently explored. Ribosome biogenesis (RB) is a critical biological process linked to various cancers; however, its systematic role in LIHC remains unclear.

METHODS: This study integrated LIHC single-cell RNA-Seq, bulk RNA-Seq, and spatial transcriptomic data with ribosome biogenesis-related gene sets to construct a single-cell atlas of LIHC. Weighted Gene Co-expression Network Analysis (WGCNA) was employed to characterize myeloid cell subsets. Furthermore, an LIHC prognostic risk model based on RB-related genes was developed using 117 machine-learning algorithm combinations. Key findings were subsequently corroborated through experimental validation and clinical sample analysis.

RESULTS: We identified a distinct macrophage subpopulation with high ribosome biogenesis activity, termed ribosome biogenesis-active macrophages (RAMs). These cells exhibited strong communication with inflammatory macrophages, potentially mediated by MIF-related receptor-ligand interactions. We further constructed an 8-gene prognostic model (PA2G4, GNL2, PWP1, DDX49, NOC4L, GDI2, CST7, and RCL1), which showed good predictive performance. Drug sensitivity analysis suggested that the high-risk group may be more responsive to several agents, including docetaxel. Among these genes, GNL2 was selected for further investigation. Elevated GNL2 expression was associated with increased stemness features in myeloid cells. Molecular docking analysis identified several candidate compounds with potential binding affinity to GNL2. Functionally, GNL2 knockdown in macrophages reduced TGF-β and TNF-α expression and was associated with decreased proliferation, migration, and invasion of LIHC cells.

CONCLUSION: We identified a highly active ribosome biogenesis-macrophage subpopulation (RAM), and constructed a robust risk model to aid in the diagnosis, prognosis, and treatment of LIHC. GNL2 is associated with increased expression of TGF-β and TNF-α and may contribute to LIHC progression.

PMID:42135716 | DOI:10.1186/s12935-026-04330-2

Mitochondrial control of glycerolipid synthesis by a PEP shuttle

SLC25A35 is revealed as the mitochondrial PEP exporter that licenses glycerolipid synthesis in lipogenic cells, linking mitochondrial energetics to lipid production. Mitochondrial PEP flux drives glycerol-3-phosphate and triglyceride synthesis, and blocking SLC25A35 reduces hepatic steatosis, offering a target for metabolic therapy.

AWPD: Frequency Shield Network for Agnostic Watermark Presence Detection

arXiv:2603.06723v2 Announce Type: replace-cross Abstract: Invisible watermarks, as an essential technology for image copyright protection, have been widely deployed with the rapid development of social media and AIGC. However, existing invisible watermark detection heavily relies on prior knowledge of specific algorithms, leading to limited detection capabilities for ``unknown watermarks'' in open environments. To this end, we propose a novel task named Agnostic Watermark Presence Detection (AWPD), which aims to identify whether an image carries a copyright mark without requiring decoding information. We construct the UniFreq-100K dataset, comprising large-scale samples across various invisible watermark embedding algorithms. Furthermore, we propose the Frequency Shield Network (FSNet). This model deploys an Adaptive Spectral Perception Module (ASPM) in the shallow layers, utilizing learnable frequency gating to dynamically amplify high-frequency watermark signals while suppressing low-frequency semantics. In the deep layers, the network introduces Dynamic Multi-Spectral Attention (DMSA) combined with tri-stream extremum pooling to deeply mine watermark energy anomalies, forcing the model to precisely focus on sensitive frequency bands. Extensive experiments demonstrate that FSNet exhibits superior zero-shot detection capabilities on the AWPD task, outperforming existing baseline models. Code and datasets will be released upon acceptance.

UWPD: A General Paradigm for Invisible Watermark Detection Agnostic to Embedding Algorithms

arXiv:2603.06723v1 Announce Type: cross Abstract: Invisible watermarks, as an essential technology for image copyright protection, have been widely deployed with the rapid development of social media and AIGC. However, existing invisible watermark detection heavily relies on prior knowledge of specific algorithms, leading to limited detection capabilities for "unknown watermarks" in open environments. To this end, we propose a novel task named Universal Watermark Presence Detection (UWPD), which aims to identify whether an image carries a copyright mark without requiring decoding information. We construct the UniFreq-100K dataset, comprising large-scale samples across various invisible watermark embedding algorithms. Furthermore, we propose the Frequency Shield Network (FSNet). This model deploys an Adaptive Spectral Perception Module (ASPM) in the shallow layers, utilizing learnable frequency gating to dynamically amplify high-frequency watermark signals while suppressing low-frequency semantics. In the deep layers, the network introduces Dynamic Multi-Spectral Attention (DMSA) combined with tri-stream extremum pooling to deeply mine watermark energy anomalies, forcing the model to precisely focus on sensitive frequency bands. Extensive experiments demonstrate that FSNet exhibits superior zero-shot detection capabilities on the UWPD task, outperforming existing baseline models. Code and datasets will be released upon acceptance.

iGVLM: Dynamic Instruction-Guided Vision Encoding for Question-Aware Multimodal Understanding

arXiv:2603.02748v2 Announce Type: replace-cross Abstract: Despite the success of Large Vision--Language Models (LVLMs), most existing architectures suffer from a representation bottleneck: they rely on static, instruction-agnostic vision encoders whose visual representations are utilized in an invariant manner across different textual tasks. This rigidity hinders fine-grained reasoning where task-specific visual cues are critical. To address this issue, we propose iGVLM, a general framework for instruction-guided visual modulation. iGVLM introduces a decoupled dual-branch architecture: a frozen representation branch that preserves task-agnostic visual representations learned during pre-training, and a dynamic conditioning branch that performs affine feature modulation via Adaptive Layer Normalization (AdaLN). This design enables a smooth transition from general-purpose perception to instruction-aware reasoning while maintaining the structural integrity and stability of pre-trained visual priors. Beyond standard benchmarks, we introduce MM4, a controlled diagnostic probe for quantifying logical consistency under multi-query, multi-instruction settings. Extensive results show that iGVLM consistently enhances instruction sensitivity across diverse language backbones, offering a plug-and-play paradigm for bridging passive perception and active reasoning.

ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion

arXiv:2603.02767v3 Announce Type: replace-cross Abstract: Image-text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organized by modality. We propose ITO, a framework addressing this limitation through two synergistic mechanisms. Multimodal multiple alignment enriches supervision by mining diverse image-text correspondences, while a lightweight training-time multimodal fusion module enforces structured cross-modal interaction. Crucially, the fusion module is discarded at inference, preserving the efficiency of standard dual-encoder architectures. Extensive experiments show that ITO consistently outperforms strong baselines across classification, retrieval, and multimodal benchmarks. Our analysis reveals that while multiple alignment drives discriminative power, training-time fusion acts as a critical structural regularizer -- eliminating the modality gap and stabilizing training dynamics to prevent the early saturation often observed in aggressive contrastive learning.

Separators in Enhancing Autoregressive Pretraining for Vision Mamba

arXiv:2603.03806v1 Announce Type: cross Abstract: The state space model Mamba has recently emerged as a promising paradigm in computer vision, attracting significant attention due to its efficient processing of long sequence tasks. Mamba's inherent causal mechanism renders it particularly suitable for autoregressive pretraining. However, current autoregressive pretraining methods are constrained to short sequence tasks, failing to fully exploit Mamba's prowess in handling extended sequences. To address this limitation, we introduce an innovative autoregressive pretraining method for Vision Mamba that substantially extends the input sequence length. We introduce new \textbf{S}epara\textbf{T}ors for \textbf{A}uto\textbf{R}egressive pretraining to demarcate and differentiate between different images, known as \textbf{STAR}. Specifically, we insert identical separators before each image to demarcate its inception. This strategy enables us to quadruple the input sequence length of Vision Mamba while preserving the original dimensions of the dataset images. Employing this long sequence pretraining technique, our STAR-B model achieved an impressive accuracy of 83.5\% on ImageNet-1k, which is highly competitive in Vision Mamba. These results underscore the potential of our method in enhancing the performance of vision models through improved leveraging of long-range dependencies.

ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion

arXiv:2603.02767v2 Announce Type: replace-cross Abstract: Image-text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organized by modality. We propose ITO, a framework addressing this limitation through two synergistic mechanisms. Multimodal multiple alignment enriches supervision by mining diverse image-text correspondences, while a lightweight training-time multimodal fusion module enforces structured cross-modal interaction. Crucially, the fusion module is discarded at inference, preserving the efficiency of standard dual-encoder architectures. Extensive experiments show that ITO consistently outperforms strong baselines across classification, retrieval, and multimodal benchmarks. Our analysis reveals that while multiple alignment drives discriminative power, training-time fusion acts as a critical structural regularizer -- eliminating the modality gap and stabilizing training dynamics to prevent the early saturation often observed in aggressive contrastive learning.

iGVLM: Dynamic Instruction-Guided Vision Encoding for Question-Aware Multimodal Understanding

arXiv:2603.02748v1 Announce Type: cross Abstract: Despite the success of Large Vision--Language Models (LVLMs), most existing architectures suffer from a representation bottleneck: they rely on static, instruction-agnostic vision encoders whose visual representations are utilized in an invariant manner across different textual tasks. This rigidity hinders fine-grained reasoning where task-specific visual cues are critical. To address this issue, we propose iGVLM, a general framework for instruction-guided visual modulation. iGVLM introduces a decoupled dual-branch architecture: a frozen representation branch that preserves task-agnostic visual representations learned during pre-training, and a dynamic conditioning branch that performs affine feature modulation via Adaptive Layer Normalization (AdaLN). This design enables a smooth transition from general-purpose perception to instruction-aware reasoning while maintaining the structural integrity and stability of pre-trained visual priors. Beyond standard benchmarks, we introduce MM4, a controlled diagnostic probe for quantifying logical consistency under multi-query, multi-instruction settings. Extensive results show that iGVLM consistently enhances instruction sensitivity across diverse language backbones, offering a plug-and-play paradigm for bridging passive perception and active reasoning.

ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion

arXiv:2603.02767v1 Announce Type: cross Abstract: Image-text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organized by modality. We propose ITO, a framework addressing this limitation through two synergistic mechanisms. Multimodal multiple alignment enriches supervision by mining diverse image-text correspondences, while a lightweight training-time multimodal fusion module enforces structured cross-modal interaction. Crucially, the fusion module is discarded at inference, preserving the efficiency of standard dual-encoder architectures. Extensive experiments show that ITO consistently outperforms strong baselines across classification, retrieval, and multimodal benchmarks. Our analysis reveals that while multiple alignment drives discriminative power, training-time fusion acts as a critical structural regularizer -- eliminating the modality gap and stabilizing training dynamics to prevent the early saturation often observed in aggressive contrastive learning.
❌