Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
Not All Tokens See Equally: Perception-Grounded Policy Optimization for Large Vision-Language Models
arXiv:2604.01840v1 Announce Type: new Abstract: While Reinforcement Learning from Verifiable Rewards (RLVR) has advanced reasoning in Large Vision-Language Models (LVLMs), prevailing frameworks suffer from a foundational methodological flaw: by distributing identical advantages across all generated tokens, these methods inherently dilute the learning signals essential for optimizing the critical, visually-grounded steps of multimodal reasoning. To bridge this gap, we formulate \textit{Token Vis
-
(Multiomics OR Omics) AND (Pancreatic)
-
Robust transcriptomic hallmarks targeting intratumor heterogeneity in intrahepatic cholangiocarcinoma
Cell Rep Med. 2026 Mar 30:102708. doi: 10.1016/j.xcrm.2026.102708. Online ahead of print.ABSTRACTIntratumor heterogeneity (ITH) undermines transcriptome-based stratification in intrahepatic cholangiocarcinoma (iCCA). Here, we integrate multi-omics data from multi-region, single-region, and single-cell RNA sequencing cohorts to systematically characterize gene expression ITH. We uncover that immune and stromal heterogeneity are primary drivers of ITH, leading to misclassification of a median 27.8
Robust transcriptomic hallmarks targeting intratumor heterogeneity in intrahepatic cholangiocarcinoma
Cell Rep Med. 2026 Mar 30:102708. doi: 10.1016/j.xcrm.2026.102708. Online ahead of print.
ABSTRACT
Intratumor heterogeneity (ITH) undermines transcriptome-based stratification in intrahepatic cholangiocarcinoma (iCCA). Here, we integrate multi-omics data from multi-region, single-region, and single-cell RNA sequencing cohorts to systematically characterize gene expression ITH. We uncover that immune and stromal heterogeneity are primary drivers of ITH, leading to misclassification of a median 27.8% of tumors by existing subtyping systems. To overcome this, we identify a low-intratumor-heterogeneity/high-intertumor-variability (LIHV) gene set and develop an ITH-insensitive classification system defining five subgroups: inflammatory (SI), metabolic (SII), atypical (SIII-1), immune-silent (SIII-2), and neurodegenerative (SIII-3). These subgroups exhibit distinct clinical outcomes, molecular features, immune landscapes, and therapeutic vulnerabilities. GPRC5A and VTCN1 serve as robust immunohistochemical biomarkers for SI and SIII tumors, while serum CEA and CA19-9 identify inflammatory iCCA. Therapeutically, HSP90 inhibition synergizes with anti-PD1 in inflammatory iCCA, whereas combined anti-PD1 and anti-TIM3 suppresses neurodegenerative iCCA. Collectively, our study provides a robust molecular framework and actionable therapeutic strategies for iCCA.
PMID:41916296 | DOI:10.1016/j.xcrm.2026.102708
-
cs.AI, q-bio.NC updates on arXiv.org
-
1S-DAug: One-Shot Data Augmentation for Robust Few-Shot Generalization
arXiv:2602.00114v3 Announce Type: replace-cross Abstract: Few-shot learning (FSL) challenges model generalization to novel classes based on just a few shots of labeled examples, a testbed where traditional test-time augmentations fail to be effective. We introduce 1S-DAug, a one-shot generative augmentation operator that synthesizes diverse yet faithful variants from just one example image at test time. 1S-DAug couples traditional geometric perturbations with controlled noise injection and a de
1S-DAug: One-Shot Data Augmentation for Robust Few-Shot Generalization
-
cs.AI, q-bio.NC updates on arXiv.org
-
MovieTeller: Tool-augmented Movie Synopsis with ID Consistent Progressive Abstraction
arXiv:2602.23228v2 Announce Type: replace-cross Abstract: With the explosive growth of digital entertainment, automated video summarization has become indispensable for applications such as content indexing, personalized recommendation, and efficient media archiving. Automatic synopsis generation for long-form videos, such as movies and TV series, presents a significant challenge for existing Vision-Language Models (VLMs). While proficient at single-image captioning, these general-purpose model
MovieTeller: Tool-augmented Movie Synopsis with ID Consistent Progressive Abstraction
-
cs.AI, q-bio.NC updates on arXiv.org
-
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
arXiv:2602.18702v1 Announce Type: cross Abstract: Long video understanding is challenging due to rich and complicated multimodal clues in long temporal range.Current methods adopt reasoning to improve the model's ability to analyze complex video clues in long videos via text-form reasoning.However,the existing literature suffers from the fact that the text-only reasoning under fixed video context may exacerbate hallucinations since detailed crucial clues are often ignored under limited video co
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
-
cs.AI, q-bio.NC updates on arXiv.org
-
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
arXiv:2602.10604v2 Announce Type: replace-cross Abstract: We introduce Step 3.5 Flash, a sparse Mixture-of-Experts (MoE) model that bridges frontier-level agentic intelligence and computational efficiency. We focus on what matters most when building agents: sharp reasoning and fast, reliable execution. Step 3.5 Flash pairs a 196B-parameter foundation with 11B active parameters for efficient inference. It is optimized with interleaved 3:1 sliding-window/full attention and Multi-Token Prediction