Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
Divide-and-Conquer Inference for Large-Scale Visual Recognition with Multimodal Large Language Models
arXiv:2605.24799v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong capabilities across a wide range of vision language tasks. However, when applied to large scale image classification, their performance degrades significantly as the label space expands a phenomenon we define as Performance Collapse in Long Sequence Recognition. Through an information theoretic analysis, we reveal that this collapse stems from a fundamental conflict between the es
-
cs.AI, q-bio.NC updates on arXiv.org
-
OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation
arXiv:2605.25829v1 Announce Type: cross Abstract: Recent vision-language-action (VLA) models and world action models (WAMs) advance robotic manipulation by enriching intermediate representations with auxiliary spatial features or future visual-state prediction. However, these representations largely remain within the observation space and do not share the rigid-body geometry of the action space, forcing the action decoder to implicitly recover this geometry. We propose OASIS, a visuomotor polic
OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook
arXiv:2604.02029v1 Announce Type: new Abstract: Latent space is rapidly emerging as a native substrate for language-based models. While modern systems are still commonly understood through explicit token-level generation, an increasing body of work shows that many critical internal processes are more naturally carried out in continuous latent space than in human-readable verbal traces. This shift is driven by the structural limitations of explicit-space computation, including linguistic redunda
The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook
-
(Multiomics OR Omics) AND (Lung OR gastric OR Hepatocellular)
-
AI-guided multi-omics analysis identifies NPC1-modulated susceptibility to SARS-CoV-2 infection under PM(2.5) exposure
Nat Commun. 2026 Mar 30. doi: 10.1038/s41467-026-71196-3. Online ahead of print.ABSTRACTExposure to airborne fine particulate matter (PM2.5) has been linked to increased risk of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection, yet the underlying mechanisms remain unclear. Here, by leveraging a fine-tuned foundation model of single-cell transcriptomics, we uncover shared transcriptional signatures between PM2.5 exposure and SARS-CoV-2 infection. We further validate this
AI-guided multi-omics analysis identifies NPC1-modulated susceptibility to SARS-CoV-2 infection under PM(2.5) exposure
Nat Commun. 2026 Mar 30. doi: 10.1038/s41467-026-71196-3. Online ahead of print.
ABSTRACT
Exposure to airborne fine particulate matter (PM2.5) has been linked to increased risk of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection, yet the underlying mechanisms remain unclear. Here, by leveraging a fine-tuned foundation model of single-cell transcriptomics, we uncover shared transcriptional signatures between PM2.5 exposure and SARS-CoV-2 infection. We further validate this association using population-level epidemiological analyses and perform genome-wide association studies (GWAS) to identify genetic variants that modulate infection risk under PM2.5 exposure. In addition, we identify NPC1 as a key modulator involved in SARS-CoV-2 infection efficiency under virus-laden PM2.5 exposure through integrative functional genomic analyses and in vitro experiments. Our findings suggest that PM2.5 facilitates viral entry through an NPC1-modulated endo-lysosomal pathway, providing a mechanistic explanation for observed pollution-related susceptibility. By integrating artificial intelligence (AI)-guided transcriptomics, epidemiology, GWAS, functional genomics, and in vitro verification, our study elucidates how environmental and genetic factors jointly influence SARS-CoV-2 susceptibility. This work highlights how AI-assisted multi-omics integration systematically decodes the health impacts of environmental exposures from molecular to population levels and informs air quality policy and infectious disease preparedness.
PMID:41912520 | DOI:10.1038/s41467-026-71196-3
-
Cell
-
Tuning the sensitivity of mechanosensory receptors through histidine scanning
Histidine scanning represents a broadly applicable technique for the identification of critical interaction sites within TCRs and other mechanosensory receptors to enhance receptor signaling strength and augment therapeutic efficacy via the catch bond mechanism.
Tuning the sensitivity of mechanosensory receptors through histidine scanning
-
cs.AI, q-bio.NC updates on arXiv.org
-
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
arXiv:2510.18632v4 Announce Type: replace-cross Abstract: Though recent advances in vision-language models (VLMs) have achieved remarkable progress across a wide range of multimodal tasks, understanding 3D spatial relationships from limited views remains a significant challenge. Previous reasoning methods typically rely on pure text (e.g., topological cognitive maps) or on 2D visual cues. However, their limited representational capacity hinders performance in specific tasks that require 3D spat
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
-
cs.AI, q-bio.NC updates on arXiv.org
-
Multimodal Continual Learning with MLLMs from Multi-scenario Perspectives
arXiv:2511.18507v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) deployed on devices must adapt to continuously changing visual scenarios such as variations in background and perspective, to effectively perform complex visual tasks. To investigate catastrophic forgetting under real-world scenario shifts, we construct a multimodal visual understanding dataset (MSVQA), covering four distinct scenarios and perspectives: high-altitude, underwater, low-altitude, and
Multimodal Continual Learning with MLLMs from Multi-scenario Perspectives
-
cs.AI, q-bio.NC updates on arXiv.org
-
Improving Multi-View Reconstruction via Texture-Guided Gaussian-Mesh Joint Optimization
arXiv:2511.03950v2 Announce Type: replace-cross Abstract: Reconstructing real-world objects from multi-view images is essential for applications in 3D editing, AR/VR, and digital content creation. Existing methods typically prioritize either geometric accuracy (Multi-View Stereo) or photorealistic rendering (Novel View Synthesis), often decoupling geometry and appearance optimization, which hinders downstream editing tasks. This paper advocates an unified treatment on geometry and appearance op
Improving Multi-View Reconstruction via Texture-Guided Gaussian-Mesh Joint Optimization
-
cs.AI, q-bio.NC updates on arXiv.org
-
GenesisGeo: Technical Report
arXiv:2509.21896v2 Announce Type: replace Abstract: Recent neuro-symbolic geometry theorem provers have made significant progress on Euclidean problems by coupling neural guidance with symbolic verification. However, most existing systems operate almost exclusively in a symbolic space, leaving diagram-based intuition largely unused during reasoning. For humans, geometric diagrams provide essential heuristics for identifying non-trivial auxiliary constructions. Meanwhile, visual language models
GenesisGeo: Technical Report
-
cs.AI, q-bio.NC updates on arXiv.org
-
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
arXiv:2602.05847v2 Announce Type: replace Abstract: While humans perceive the world through diverse modalities that operate synergistically to support a holistic understanding of their surroundings, existing omnivideo models still face substantial challenges on audio-visual understanding tasks. In this paper, we propose OmniVideo-R1, a novel reinforced framework that improves mixed-modality reasoning. OmniVideo-R1 empowers models to "think with omnimodal cues" by two key strategies: (1) query-i
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
-
cs.AI, q-bio.NC updates on arXiv.org
-
AECBench: A Hierarchical Benchmark for Knowledge Evaluation of Large Language Models in the AEC Field
arXiv:2509.18776v3 Announce Type: replace-cross Abstract: Large language models (LLMs), as a novel information technology, are seeing increasing adoption in the Architecture, Engineering, and Construction (AEC) field. They have shown their potential to streamline processes throughout the building lifecycle. However, the robustness and reliability of LLMs in such a specialized and safety-critical domain remain to be evaluated. To address this challenge, this paper establishes AECBench, a compreh