❌

Normal view

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

arXiv:2607.19363v3 Announce Type: replace Abstract: Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads. Using simplified retrieval tasks and length generalization scenarios, we show -- both empirically and theoretically -- that heads with different functional roles require distinct frequency ranges and attention scaling factors to operate effectively. Ignoring this structure leads to suboptimal utilization of embedding dimensions and degraded performance, particularly under long-context settings. To address these limitations, we propose AdaRoPE, which equips each attention head with learnable rotation frequencies and attention scaling factors. Pretrained LLMs with AdaRoPE consistently outperform existing RoPE variants, including partial RoPE and NoPE baselines. For context extension, we further show that uniform frequency and attention scaling, used in methods such as YaRN, are suboptimal. By applying head-specific scaling, AdaRoPE enables better context extension while better preserving short-context performance in both the extrapolation setting and the long-context continued pretraining setting. These results highlight the importance of optimizing rotary position embedding at the level of individual attention heads.

Flow-OPD: On-Policy Distillation for Flow Matching Models

arXiv:2605.08063v5 Announce Type: replace-cross Abstract: Existing Flow Matching (FM) text-to-image models suffer from two critical bottlenecks under multi-task alignment: the reward sparsity induced by scalar-valued rewards, and the gradient interference arising from jointly optimizing heterogeneous objectives, which together give rise to a 'seesaw effect' of competing metrics and pervasive reward hacking. Inspired by the success of On-Policy Distillation (OPD) in the large language model community, we propose Flow-OPD, the first unified post-training framework that integrates on-policy distillation into Flow Matching models. Flow-OPD adopts a two-stage alignment strategy: it first cultivates domain-specialized teacher models via single-reward GRPO fine-tuning, allowing each expert to reach its performance ceiling in isolation; it then establishes a robust initial policy through a Flow-based Cold-Start scheme and seamlessly consolidates heterogeneous expertise into a single student via a three-step orchestration of on-policy sampling, task-routing labeling, and dense trajectory-level supervision. We further introduce Manifold Anchor Regularization (MAR), which leverages a task-agnostic teacher to provide full-data supervision that anchors generation to a high-quality manifold, effectively mitigating the aesthetic degradation commonly observed in purely RL-driven alignment. Built upon Stable Diffusion 3.5 Medium, Flow-OPD raises the GenEval score from 63 to 92 and the OCR accuracy from 59 to 94, yielding an overall improvement of roughly 10 points over vanilla GRPO, while preserving image fidelity and human-preference alignment and exhibiting an emergent 'teacher-surpassing' effect. These results establish Flow-OPD as a scalable alignment paradigm for building generalist text-to-image models. The codes and weights will be released in: https://github.com/CostaliyA/Flow-OPD .

Establishment and characterization of an immortalized porcine gastric epithelial cell line and identification of NPC1 as a key mediator of aflatoxin B1 toxicity

Gene. 2026 Apr 9:150160. doi: 10.1016/j.gene.2026.150160. Online ahead of print.

ABSTRACT

Porcine gastric epithelial cells (PGECs) serve as a valuable model for studying the molecular and pathogenic mechanisms of the stomach. However, PGECs face limitations such as isolation challenges, short lifespan, and restricted proliferation. To address this, we established an immortalized PGECs (i-PGECs) to enable in vitro investigation of pathogen infection mechanisms. Primary PGECs were isolated from the acid-secreting glands using stepwise digestion with multiple enzymes (dispase II/collagenase I/hyaluronidase). Immortalization was achieved via lentiviral vectors expressing simian virus 40 large T antigen (SV40T) and human telomerase reverse transcriptase (hTERT), with successful expression confirmed by qRT-PCR (P < 0.05). Epithelial identity of i-PGECs was confirmed by stable expression of CK18, EpCAM, and E-cadherin, as shown by qRT-PCR and immunofluorescence. i-PGECs retained the morphological and ultrastructural features of PGECs and exhibited enhanced proliferation, as demonstrated by WST-8 assays, apoptosis and cell cycle analysis, karyotyping, and transmission electron microscopy (TEM). Telomere length analysis and scratch wound assays demonstrated stable telomere maintenance and consistent migration capacity unaffected by passaging. RNA-sequencing and differential expressed genes (DEGs) analysis revealed significantly upregulating of genes involved in cell proliferation pathways (P < 0.01). Following aflatoxin B1 (AFB1) exposure, i-PGECs significantly upregulated immune-related factors, such as NPC1 and PLAUR (P < 0.01). CRISPR/Cas9-mediated knockout of NPC1 in i-PGECs conferred increased resistance to AFB1-induced cytotoxicity, as shown by WST-8 assay. The i-PGECs remained stable after more than 50 passages, supporting their use as a reliable for in vitro model investigating the mechanisms of toxicity infection in the porcine gastric epithelium.

PMID:41966285 | DOI:10.1016/j.gene.2026.150160

Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models

arXiv:2601.22060v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved remarkable success across a broad range of vision tasks. However, constrained by the capacity of their internal world knowledge, prior work has proposed augmenting MLLMs by ``reasoning-then-tool-call'' for visual and textual search engines to obtain substantial gains on tasks requiring extensive factual information. However, these approaches typically define multimodal search in a naive setting, assuming that a single full-level or entity-level image query and few text query suffices to retrieve the key evidence needed to answer the question, which is unrealistic in real-world scenarios with substantial visual noise. Moreover, they are often limited in the reasoning depth and search breadth, making it difficult to solve complex questions that require aggregating evidence from diverse visual and textual sources. Building on this, we propose Vision-DeepResearch, which proposes one new multimodal deep-research paradigm, i.e., performs multi-turn, multi-entity and multi-scale visual and textual search to robustly hit real-world search engines under heavy noise. Our Vision-DeepResearch supports dozens of reasoning steps and hundreds of engine interactions, while internalizing deep-research capabilities into the MLLM via cold-start supervision and RL training, resulting in a strong end-to-end multimodal deep-research MLLM. It substantially outperforming existing multimodal deep-research MLLMs, and workflows built on strong closed-source foundation model such as GPT-5, Gemini-2.5-pro and Claude-4-Sonnet. The code will be released in https://github.com/Osilly/Vision-DeepResearch.
❌