Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
arXiv:2609.09113v2 Announce Type: replace Abstract: While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essential to bridge this gap, among which Sparse Autoencoders (SAEs) serve as a cornerstone by isolating interpretable features for model inspection and
-
Molecular Therapy
-
Targeting the MNK1-MYH9 axis blocks YAP1 recruitment to prevent thrombosis and platelet activation-induced NETosis
MNK1 acts as a structural shield on MYH9, preventing YAP1-mediated platelet activation. Developing MD2 to lock this MNK1-MYH9 complex introduces a safe antithrombotic strategy, shifting the therapeutic paradigm from kinase inhibition to stabilizing protein-protein interactions against immunothrombosis.
Targeting the MNK1-MYH9 axis blocks YAP1 recruitment to prevent thrombosis and platelet activation-induced NETosis
-
cs.AI, q-bio.NC updates on arXiv.org
-
RAU: Reference-based Anatomical Understanding with Vision Language Models
arXiv:2509.22404v2 Announce Type: replace-cross Abstract: Anatomical understanding, which is the ability to identify, localize, or segment anatomical structures, is critical in medical image analysis; however, its progress is constrained by the scarcity of expert-labeled data. A promising remedy is to leverage an annotated reference image to guide the interpretation of an unlabeled target. Although recent vision-language models (VLMs) exhibit non-trivial visual reasoning, their reference-based
RAU: Reference-based Anatomical Understanding with Vision Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization
arXiv:2605.10764v2 Announce Type: replace-cross Abstract: Recent studies show that gradient-based universal image jailbreaks on vision-language models (VLMs) exhibit little or no cross-model transferability, casting doubt on the feasibility of transferable multimodal jailbreaks. We revisit this conclusion under a strictly untargeted threat model without enforcing a fixed prefix or response pattern. Our preliminary experiment reveals that refusal behavior concentrates at high-entropy tokens duri
Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization
-
Omics in Gastric
-
FDX1 as a predictive biomarker and therapeutic target for lymph node metastasis in gastric cancer
Clin Exp Med. 2026 May 10. doi: 10.1007/s10238-026-02160-0. Online ahead of print.ABSTRACTThe prognostic values of cuproptosis-related genes (CRGs) in gastric cancer with lymph node metastasis (GCLM), especially in the tumor immune microenvironment (TIME), remain unclear. We analyzed the expression, mutation, immunity, drug sensitivity, and prognostic value of CRGs in GCLM using TCGA and GEO cohorts. Consensus clustering was performed to identify CRG subtypes, with differences characterized by m
FDX1 as a predictive biomarker and therapeutic target for lymph node metastasis in gastric cancer
Clin Exp Med. 2026 May 10. doi: 10.1007/s10238-026-02160-0. Online ahead of print.
ABSTRACT
The prognostic values of cuproptosis-related genes (CRGs) in gastric cancer with lymph node metastasis (GCLM), especially in the tumor immune microenvironment (TIME), remain unclear. We analyzed the expression, mutation, immunity, drug sensitivity, and prognostic value of CRGs in GCLM using TCGA and GEO cohorts. Consensus clustering was performed to identify CRG subtypes, with differences characterized by multi-omics analysis. A CRG-based prognostic risk score and immune score were constructed for individualized assessment, and the role of CRGs was validated through in vitro and in vivo experiments. Consensus clustering revealed that CRGs were significantly enriched in biological processes related to mitosis and energy metabolism, as well as in immune-related and cancer-associated pathways. Four distinct CRG subtypes were identified, showing marked differences in expression profiles, prognosis, genetic alterations, TIME, and chemotherapeutic drug sensitivity. We developed an exploratory CRG-based prognostic risk score for preliminary individualized assessment, and the functional relevance of CRGs in GCLM was further validated through in vitro experiments. Among these, FDX1, LIAS, DLAT, MTF1, and GLS were identified as key determinants of overall survival in patients with GCLM, with FDX1 emerging as a potential independent prognostic factor. Notably, upregulation of FDX1 significantly suppressed lymph node metastasis of gastric cancer cells in a mouse popliteal lymph node metastasis model. Our data uncovers FDX1 might be a potential favorable prognostic factors in GCLM patients. These findings may improve our understanding of CRGs in GCLM and provide new in-sights for assessing prognosis and developing more effective treatment strategies.
PMID:42107026 | DOI:10.1007/s10238-026-02160-0
-
Nature - Issue - nature.com science feeds
-
Composable neural emulators accelerate thermoelectric generator design
Nature, Published online: 15 April 2026; doi:10.1038/s41586-026-10223-1A composable neural network emulator is described for speeding up thermoelectric generator design, demonstrating the ability to predict generator performance with >99% accuracy while taking only 0.01% of the time compared with commercial finite-element solvers.
Composable neural emulators accelerate thermoelectric generator design
Nature, Published online: 15 April 2026; doi:10.1038/s41586-026-10223-1
A composable neural network emulator is described for speeding up thermoelectric generator design, demonstrating the ability to predict generator performance with >99% accuracy while taking only 0.01% of the time compared with commercial finite-element solvers.-
cs.AI, q-bio.NC updates on arXiv.org
-
ResAdapt: Adaptive Resolution for Efficient Multimodal Reasoning
arXiv:2603.28610v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) achieve stronger visual understanding by scaling input fidelity, yet the resulting visual token growth makes jointly sustaining high spatial resolution and long temporal context prohibitive. We argue that the bottleneck lies not in how post-encoding representations are compressed but in the volume of pixels the encoder receives, and address it with ResAdapt, an Input-side adaptation framework that
ResAdapt: Adaptive Resolution for Efficient Multimodal Reasoning
-
cs.AI, q-bio.NC updates on arXiv.org
-
From Editor to Dense Geometry Estimator
arXiv:2509.04338v2 Announce Type: replace-cross Abstract: Leveraging visual priors from pre-trained text-to-image (T2I) generative models has shown success in dense prediction. However, dense prediction is inherently an image-to-image task, suggesting that image editing models, rather than T2I generative models, may be a more suitable foundation for fine-tuning. Motivated by this, we conduct a systematic analysis of the fine-tuning behaviors of both editors and generators for dense geometry e
From Editor to Dense Geometry Estimator
-
cs.AI, q-bio.NC updates on arXiv.org
-
Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench
arXiv:2510.26865v2 Announce Type: replace-cross Abstract: Reading measurement instruments is effortless for humans and requires relatively little domain expertise, yet it remains surprisingly challenging for current vision-language models (VLMs) as we find in preliminary evaluation. In this work, we introduce MeasureBench, a benchmark on visual measurement reading covering both real-world and synthesized images of various types of measurements, along with an extensible pipeline for data synthes
Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench
-
cs.AI, q-bio.NC updates on arXiv.org
-
Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
arXiv:2512.03454v3 Announce Type: replace-cross Abstract: Interpreting natural-language commands to localize target objects is critical for autonomous driving (AD). Existing visual grounding (VG) methods for autonomous vehicles (AVs) typically struggle with ambiguous, context-dependent instructions, as they lack reasoning over 3D spatial relations and anticipated scene evolution. Grounded in the principles of world models, we propose ThinkDeeper, a framework that reasons about future spatial st
Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
-
cs.AI, q-bio.NC updates on arXiv.org
-
Thinking in Streaming Video
arXiv:2603.12938v1 Announce Type: cross Abstract: Real-time understanding of continuous video streams is essential for interactive assistants and multimodal agents operating in dynamic environments. However, most existing video reasoning approaches follow a batch paradigm that defers reasoning until the full video context is observed, resulting in high latency and growing computational cost that are incompatible with streaming scenarios. In this paper, we introduce ThinkStream, a framework for
Thinking in Streaming Video
-
cs.AI, q-bio.NC updates on arXiv.org
-
Evo: Autoregressive-Diffusion Large Language Models with Evolving Balance
arXiv:2603.06617v1 Announce Type: cross Abstract: We introduce \textbf{Evo}, a duality latent trajectory model that bridges autoregressive (AR) and diffusion-based language generation within a continuous evolutionary generative framework. Rather than treating AR decoding and diffusion generation as separate paradigms, Evo reconceptualizes text generation as a latent flow: each token is associated with a vector-valued embedding that evolves over a progression variable $t_i \in [0, 1]$, indicatin
Evo: Autoregressive-Diffusion Large Language Models with Evolving Balance
-
cs.AI, q-bio.NC updates on arXiv.org
-
Multifaceted Scenario-Aware Hypergraph Learning for Next POI Recommendation
arXiv:2601.11610v2 Announce Type: replace-cross Abstract: Among the diverse services provided by Location-Based Social Networks (LBSNs), Next Point-of-Interest (POI) recommendation plays a crucial role in inferring user preferences from historical check-in trajectories. However, existing sequential and graph-based methods frequently neglect significant mobility variations across distinct contextual scenarios (e.g., tourists versus locals). This oversight results in suboptimal performance due to
Multifaceted Scenario-Aware Hypergraph Learning for Next POI Recommendation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Geometry-Guided Reinforcement Learning for Multi-view Consistent 3D Scene Editing
arXiv:2603.03143v1 Announce Type: cross Abstract: Leveraging the priors of 2D diffusion models for 3D editing has emerged as a promising paradigm. However, maintaining multi-view consistency in edited results remains challenging, and the extreme scarcity of 3D-consistent editing paired data renders supervised fine-tuning (SFT), the most effective training strategy for editing tasks, infeasible. In this paper, we observe that, while generating multi-view consistent 3D content is highly challengi
Geometry-Guided Reinforcement Learning for Multi-view Consistent 3D Scene Editing
-
cs.AI, q-bio.NC updates on arXiv.org
-
VLANeXt: Recipes for Building Strong VLA Models
arXiv:2602.18532v1 Announce Type: cross Abstract: Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding for general-purpose policy learning. Yet, the current VLA landscape remains fragmented and exploratory. Although many groups have proposed their own VLA models, inconsistencies in training protocols and evaluation settings make it difficult to identify which design choices truly matter. To bring structu
VLANeXt: Recipes for Building Strong VLA Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
HONEST-CAV: Hierarchical Optimization of Network Signals and Trajectories for Connected and Automated Vehicles with Multi-Agent Reinforcement Learning
arXiv:2602.18740v1 Announce Type: cross Abstract: This study presents a hierarchical, network-level traffic flow control framework for mixed traffic consisting of Human-driven Vehicles (HVs), Connected and Automated Vehicles (CAVs). The framework jointly optimizes vehicle-level eco-driving behaviors and intersection-level traffic signal control to enhance overall network efficiency and decrease energy consumption. A decentralized Multi-Agent Reinforcement Learning (MARL) approach by Value Decom
HONEST-CAV: Hierarchical Optimization of Network Signals and Trajectories for Connected and Automated Vehicles with Multi-Agent Reinforcement Learning
-
cs.AI, q-bio.NC updates on arXiv.org
-
Taming Preconditioner Drift: Unlocking the Potential of Second-Order Optimizers for Federated Learning on Non-IID Data
arXiv:2602.19271v1 Announce Type: cross Abstract: Second-order optimizers can significantly accelerate large-scale training, yet their naive federated variants are often unstable or even diverge on non-IID data. We show that a key culprit is \emph{preconditioner drift}: client-side second-order training induces heterogeneous \emph{curvature-defined geometries} (i.e., preconditioner coordinate systems), and server-side model averaging updates computed under incompatible metrics, corrupting the
Taming Preconditioner Drift: Unlocking the Potential of Second-Order Optimizers for Federated Learning on Non-IID Data
-
cs.AI, q-bio.NC updates on arXiv.org
-
Rethinking LoRA for Privacy-Preserving Federated Learning in Large Models
arXiv:2602.19926v1 Announce Type: cross Abstract: Fine-tuning large vision models (LVMs) and large language models (LLMs) under differentially private federated learning (DPFL) is hindered by a fundamental privacy-utility trade-off. Low-Rank Adaptation (LoRA), a promising parameter-efficient fine-tuning (PEFT) method, reduces computational and communication costs by introducing two trainable low-rank matrices while freezing pre-trained weights. However, directly applying LoRA in DPFL settings l
Rethinking LoRA for Privacy-Preserving Federated Learning in Large Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
DP-FedAdamW: An Efficient Optimizer for Differentially Private Federated Large Models
arXiv:2602.19945v1 Announce Type: cross Abstract: Balancing convergence efficiency and robustness under Differential Privacy (DP) is a central challenge in Federated Learning (FL). While AdamW accelerates training and fine-tuning in large-scale models, we find that directly applying it to Differentially Private FL (DPFL) suffers from three major issues: (i) data heterogeneity and privacy noise jointly amplify the variance of second-moment estimator, (ii) DP perturbations bias the second-moment
DP-FedAdamW: An Efficient Optimizer for Differentially Private Federated Large Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
arXiv:2508.05612v5 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as an effective post-training paradigm for enhancing the reasoning capabilities of multimodal large language model (MLLM). However, current RL pipelines often suffer from training inefficiencies caused by two underexplored issues: Advantage Collapsing, where most advantages in a batch concentrate near zero, and Rollout Silencing, where the proportion of rollouts contributing non-zero gradients dimi