Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
LR-SGS: Robust LiDAR-Reflectance-Guided Salient Gaussian Splatting for Self-Driving Scene Reconstruction
arXiv:2603.12647v1 Announce Type: cross Abstract: Recent 3D Gaussian Splatting (3DGS) methods have demonstrated the feasibility of self-driving scene reconstruction and novel view synthesis. However, most existing methods either rely solely on cameras or use LiDAR only for Gaussian initialization or depth supervision, while the rich scene information contained in point clouds, such as reflectance, and the complementarity between LiDAR and RGB have not been fully exploited, leading to degradatio
-
cs.AI, q-bio.NC updates on arXiv.org
-
FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restoration
arXiv:2511.14099v3 Announce Type: replace-cross Abstract: All-in-One Image Restoration (AIO-IR) aims to develop a unified model that can handle multiple degradations under complex conditions. However, existing methods often rely on task-specific designs or latent routing strategies, making it hard to adapt to real-world scenarios with various degradations. We propose FAPE-IR, a Frequency-Aware Planning and Execution framework for image restoration. It uses a frozen Multimodal Large Language Mod
FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restoration
-
cs.AI, q-bio.NC updates on arXiv.org
-
Multimodal Continual Learning with MLLMs from Multi-scenario Perspectives
arXiv:2511.18507v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) deployed on devices must adapt to continuously changing visual scenarios such as variations in background and perspective, to effectively perform complex visual tasks. To investigate catastrophic forgetting under real-world scenario shifts, we construct a multimodal visual understanding dataset (MSVQA), covering four distinct scenarios and perspectives: high-altitude, underwater, low-altitude, and
Multimodal Continual Learning with MLLMs from Multi-scenario Perspectives
-
Omics in Gastric
-
Targeting the OXNAD1-PTGS2 axis with resveratrol overcomes ferroptosis Inhibition and reverses 5-FU resistance in gastric cancer
Gastric Cancer. 2026 Mar 13. doi: 10.1007/s10120-026-01718-x. Online ahead of print.ABSTRACTBACKGROUND: 5-Fluorouracil (5-FU) remains a cornerstone of first-line chemotherapy for gastric cancer, yet the emergence of resistance severely compromises its clinical efficacy. Although ferroptosis suppression has been recognized as a pivotal mechanism of chemoresistance, the mitochondrial regulatory processes involved remain poorly understood.METHODS: We integrated clinical specimen analysis, in vitro
Targeting the OXNAD1-PTGS2 axis with resveratrol overcomes ferroptosis Inhibition and reverses 5-FU resistance in gastric cancer
Gastric Cancer. 2026 Mar 13. doi: 10.1007/s10120-026-01718-x. Online ahead of print.
ABSTRACT
BACKGROUND: 5-Fluorouracil (5-FU) remains a cornerstone of first-line chemotherapy for gastric cancer, yet the emergence of resistance severely compromises its clinical efficacy. Although ferroptosis suppression has been recognized as a pivotal mechanism of chemoresistance, the mitochondrial regulatory processes involved remain poorly understood.
METHODS: We integrated clinical specimen analysis, in vitro and in vivo functional assays, multi-omics profiling, and molecular docking to delineate the role of the mitochondrial oxidoreductase OXNAD1 in mediating 5-FU resistance in gastric cancer, and to assess the therapeutic potential of the natural polyphenol resveratrol as a chemosensitizing agent.
RESULTS: OXNAD1 was found to be significantly overexpressed in gastric cancer tissues and cell lines, correlating with unfavorable prognosis and enhanced 5-FU resistance. Mechanistically, OXNAD1 directly bound to and suppressed the ferroptosis driver PTGS2, thereby attenuating lipid peroxidation and mitochondrial damage, ultimately restraining ferroptosis and promoting drug resistance. Notably, resveratrol disrupted the OXNAD1-PTGS2 interaction by directly binding OXNAD1, reinstating ferroptotic activity, markedly enhancing the cytotoxic effect of 5-FU in resistant cells, and potentiating the antitumor efficacy of 5-FU in xenograft models.
CONCLUSION: The OXNAD1-PTGS2 axis constitutes a critical metabolic-cell death cross-regulatory pathway underlying 5-FU resistance in gastric cancer. Targeting this axis with resveratrol provides a promising combinatorial strategy to overcome chemoresistance.
PMID:41824193 | DOI:10.1007/s10120-026-01718-x
-
cs.AI, q-bio.NC updates on arXiv.org
-
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
arXiv:2603.07540v1 Announce Type: cross Abstract: Unified multimodal models hold the promise of generating extensive, interleaved narratives, weaving text and imagery into coherent long-form stories. However, current systems suffer from a critical reliability gap: as sequences grow, generation quality rapidly collapses. In this work, we investigate the mechanism behind this failure and argue that it is distinct from standard long-context challenges. We reveal that in generation, accumulated vis
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
arXiv:2508.09486v2 Announce Type: replace-cross Abstract: Video Large Language Models (Video-LLMs) have shown strong video understanding, yet their application to long-form videos remains constrained by limited context windows. A common workaround is to compress long videos into a handful of representative frames via retrieval or summarization. However, most existing pipelines score frames in isolation, implicitly assuming that frame-level saliency is sufficient for downstream reasoning. This o
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
-
cs.AI, q-bio.NC updates on arXiv.org
-
LMMRec: LLM-driven Motivation-aware Multimodal Recommendation
arXiv:2602.05474v3 Announce Type: replace-cross Abstract: Motivation-based recommendation systems uncover user behavior drivers. Motivation modeling, crucial for decision-making and content preference, explains recommendation generation. Existing methods often treat motivation as latent variables from interaction data, neglecting heterogeneous information like review text. In multimodal motivation fusion, two challenges arise: 1) achieving stable cross-modal alignment amid noise, and 2) identif
LMMRec: LLM-driven Motivation-aware Multimodal Recommendation
-
cs.AI, q-bio.NC updates on arXiv.org
-
AI4S-SDS: A Neuro-Symbolic Solvent Design System via Sparse MCTS and Differentiable Physics Alignment
arXiv:2603.03686v1 Announce Type: new Abstract: Automated design of chemical formulations is a cornerstone of materials science, yet it requires navigating a high-dimensional combinatorial space involving discrete compositional choices and continuous geometric constraints. Existing Large Language Model (LLM) agents face significant challenges in this setting, including context window limitations during long-horizon reasoning and path-dependent exploration that may lead to mode collapse. To addr
AI4S-SDS: A Neuro-Symbolic Solvent Design System via Sparse MCTS and Differentiable Physics Alignment
-
cs.AI, q-bio.NC updates on arXiv.org
-
Specification-Driven Generation and Evaluation of Discrete-Event World Models via the DEVS Formalism
arXiv:2603.03784v1 Announce Type: new Abstract: World models are essential for planning and evaluation in agentic systems, yet existing approaches lie at two extremes: hand-engineered simulators that offer consistency and reproducibility but are costly to adapt, and implicit neural models that are flexible but difficult to constrain, verify, and debug over long horizons. We seek a principled middle ground that combines the reliability of explicit simulators with the flexibility of learned model
Specification-Driven Generation and Evaluation of Discrete-Event World Models via the DEVS Formalism
-
cs.AI, q-bio.NC updates on arXiv.org
-
UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?
arXiv:2603.03241v1 Announce Type: cross Abstract: Unified multimodal models have recently demonstrated strong generative capabilities, yet whether and when generation improves understanding remains unclear. Existing benchmarks lack a systematic exploration of the specific tasks where generation facilitates understanding. To this end, we introduce UniG2U-Bench, a comprehensive benchmark categorizing generation-to-understanding (G2U) evaluation into 7 regimes and 30 subtasks, requiring varying de
UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?
-
cs.AI, q-bio.NC updates on arXiv.org
-
Leverage Knowledge Graph and Large Language Model for Law Article Recommendation: A Case Study of Chinese Criminal Law
arXiv:2410.04949v3 Announce Type: replace-cross Abstract: Judicial efficiency is critical to social stability. However, in many countries worldwide, grassroots courts face substantial case backlogs, and judicial decisions remain heavily dependent on judges' cognitive efforts, with insufficient intelligent tools to enhance efficiency. To address this issue, we propose a highly efficient law article recommendation approach combining a Knowledge Graph (KG) and a Large Language Model (LLM). First,
Leverage Knowledge Graph and Large Language Model for Law Article Recommendation: A Case Study of Chinese Criminal Law
-
cs.AI, q-bio.NC updates on arXiv.org
-
xLLM Technical Report
arXiv:2510.14686v2 Announce Type: replace-cross Abstract: We introduce xLLM, an intelligent and efficient Large Language Model (LLM) inference framework designed for high-performance, large-scale enterprise-grade serving, with deep optimizations for diverse AI accelerators. To address these challenges, xLLM builds a novel decoupled service-engine architecture. At the service layer, xLLM-Service features an intelligent scheduling module that efficiently processes multimodal requests and co-locat
xLLM Technical Report
-
cs.AI, q-bio.NC updates on arXiv.org
-
DeepXiv-SDK: An Agentic Data Interface for Scientific Literature
arXiv:2603.00084v2 Announce Type: replace-cross Abstract: LLM-agents are increasingly used to accelerate the progress of scientific research. Yet a persistent bottleneck is data access: agents not only lack readily available tools for retrieval, but also have to work with unstrcutured, human-centric data on the Internet, such as HTML web-pages and PDF files, leading to excessive token consumption, limit working efficiency, and brittle evidence look-up. This gap motivates the development of \tex
DeepXiv-SDK: An Agentic Data Interface for Scientific Literature
-
cs.AI, q-bio.NC updates on arXiv.org
-
Efficient endometrial carcinoma screening via cross-modal synthesis and gradient distillation
arXiv:2602.19822v1 Announce Type: cross Abstract: Early detection of myometrial invasion is critical for the staging and life-saving management of endometrial carcinoma (EC), a prevalent global malignancy. Transvaginal ultrasound serves as the primary, accessible screening modality in resource-constrained primary care settings; however, its diagnostic reliability is severely hindered by low tissue contrast, high operator dependence, and a pronounced scarcity of positive pathological samples. Ex
Efficient endometrial carcinoma screening via cross-modal synthesis and gradient distillation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Diversity-Incentivized Exploration for Versatile Reasoning
arXiv:2509.26209v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a crucial paradigm for incentivizing reasoning capabilities in Large Language Models (LLMs). Due to vast state-action spaces and reward sparsity in reasoning tasks, existing methods often struggle with deficient exploration and poor sample efficiency. In the paper, we propose \textbf{DIVER} (\textbf{D}iversity-\textbf{I}ncentivized Exploration for \textbf{V}ersatil\textbf{E}
Diversity-Incentivized Exploration for Versatile Reasoning
-
cs.AI, q-bio.NC updates on arXiv.org
-
Dialogue is Better Than Monologue: Instructing Medical LLMs via Strategical Conversations
arXiv:2501.17860v2 Announce Type: replace-cross Abstract: Current medical AI systems often fail to replicate real-world clinical reasoning, as they are predominantly trained and evaluated on static text and question-answer tasks. These tuning methods and benchmarks overlook critical aspects like evidence-based reasoning and handling distracting information. To bridge this gap, we introduce a novel benchmark that simulates real-world diagnostic scenarios, integrating noise and difficulty levels
Dialogue is Better Than Monologue: Instructing Medical LLMs via Strategical Conversations
-
cs.AI, q-bio.NC updates on arXiv.org
-
STaRR: Spatial-Temporal Token-Dynamics-Aware Responsive Remasking for Diffusion Language Models
arXiv:2601.04205v2 Announce Type: replace-cross Abstract: Diffusion Language Models (DLMs) enable parallel decoding via iterative denoising, where remasking strategies play a critical role in balancing inference speed and output quality. Existing methods predominantly rely on static confidence thresholds, overlooking the spatial-temporal dynamics of token confidence, causing unnecessary remasking. We propose Spatial-Temporal Token-Dynamics-Aware Responsive Remasking (STaRR), a training-free fra
STaRR: Spatial-Temporal Token-Dynamics-Aware Responsive Remasking for Diffusion Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
arXiv:2602.10604v2 Announce Type: replace-cross Abstract: We introduce Step 3.5 Flash, a sparse Mixture-of-Experts (MoE) model that bridges frontier-level agentic intelligence and computational efficiency. We focus on what matters most when building agents: sharp reasoning and fast, reliable execution. Step 3.5 Flash pairs a 196B-parameter foundation with 11B active parameters for efficient inference. It is optimized with interleaved 3:1 sliding-window/full attention and Multi-Token Prediction
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
arXiv:2602.14364v1 Announce Type: cross Abstract: Clawdbot is a self-hosted, tool-using personal AI agent with a broad action space spanning local execution and web-mediated workflows, which raises heightened safety and security concerns under ambiguity and adversarial steering. We present a trajectory-centric evaluation of Clawdbot across six risk dimensions. Our test suite samples and lightly adapts scenarios from prior agent-safety benchmarks (including ATBench and LPS-Bench) and supplements
A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
-
cs.AI, q-bio.NC updates on arXiv.org
-
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
arXiv:2510.10689v2 Announce Type: replace Abstract: Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modalities or integrating them in a logically inconsistent manner. To bridge this gap, we introduce OmniVideoBench, a large-scale and rigorously designed b