Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
arXiv:2603.23184v1 Announce Type: cross Abstract: Reward modeling represents a long-standing challenge in reinforcement learning from human feedback (RLHF) for aligning language models. Current reward modeling is heavily contingent upon experimental feedback data with high collection costs. In this work, we study \textit{implicit reward modeling} -- learning reward models from implicit human feedback (e.g., clicks and copies) -- as a cost-effective alternative. We identify two fundamental chall
-
cs.AI, q-bio.NC updates on arXiv.org
-
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
arXiv:2601.13719v2 Announce Type: replace-cross Abstract: Long video understanding presents significant challenges for vision-language models due to extremely long context windows. Existing solutions relying on naive chunking strategies with retrieval-augmented generation, typically suffer from information fragmentation and a loss of global coherence. We present HAVEN, a unified framework for long-video understanding that enables coherent and comprehensive reasoning by integrating audiovisual e
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
-
cs.AI, q-bio.NC updates on arXiv.org
-
When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning
arXiv:2603.21289v2 Announce Type: replace-cross Abstract: Recent progress in multimodal large language models has led to strong performance on reasoning tasks, but these improvements largely rely on high-quality annotated data or teacher-model distillation, both of which are costly and difficult to scale. To address this, we propose an unsupervised self-evolution training framework for multimodal reasoning that achieves stable performance improvements without using human-annotated answers or ex
When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning
-
Omics in Hepatocellular
-
SIRT3 deacetylates STEAP4 to modulate cuproptosis sensitivity via mitochondrial metabolic reprogramming in HBV-related HCC
Cell Death Differ. 2026 Mar 16. doi: 10.1038/s41418-026-01713-w. Online ahead of print.ABSTRACTHepatitis B virus (HBV) infection remains a leading etiological driver of hepatocellular carcinoma (HCC). Cuproptosis is a recently defined copper-dependent form of regulated cell death that selectively eliminates mitochondria-dependent cells; whether HBV rewires this vulnerability remains unknown. Here we unveil a novel HBV X protein (HBx)-driven mechanism of cuproptosis evasion. Integrative analysis
SIRT3 deacetylates STEAP4 to modulate cuproptosis sensitivity via mitochondrial metabolic reprogramming in HBV-related HCC
Cell Death Differ. 2026 Mar 16. doi: 10.1038/s41418-026-01713-w. Online ahead of print.
ABSTRACT
Hepatitis B virus (HBV) infection remains a leading etiological driver of hepatocellular carcinoma (HCC). Cuproptosis is a recently defined copper-dependent form of regulated cell death that selectively eliminates mitochondria-dependent cells; whether HBV rewires this vulnerability remains unknown. Here we unveil a novel HBV X protein (HBx)-driven mechanism of cuproptosis evasion. Integrative analysis of clinical specimens, HBx-transgenic (HBx-Tg) mice, and multi-omics datasets revealed marked downregulation of STEAP4 (six-transmembrane epithelial antigen of prostate 4), a metalloreductase essential for cuproptosis sensitivity, in HBV-positive HCC. Mechanistically, HBx attenuates sirtuin 3 (SIRT3), impairing deacetylation of STEAP4 at lysine 404 and abolishing its mitochondrial targeting. Consequently, cells switch from the tricarboxylic acid (TCA) cycle respiration to glycolysis, reducing sensitivity to the copper ionophore elesclomol (ES). Restoring STEAP4 expression or pharmacological activation of SIRT3 with honokiol (HKL) re-instated mitochondrial STEAP4 localization and re-sensitized HBV-related HCC cells to cuproptosis; combination with ES produced synergistic tumor suppression in vitro and in orthotopic models. Collectively, our findings establish the SIRT3-STEAP4 axis as a novel regulator of cuproptosis resistance in HBV-related HCC. HBx-mediated repression of SIRT3 disrupts STEAP4 deacetylation and mitochondrial targeting, fostering metabolic reprogramming and evasion of copper-induced cell death. The results provide a pre-clinical rationale for copper-directed combination strategies in HBV-associated HCC.
PMID:41840161 | DOI:10.1038/s41418-026-01713-w
-
cs.AI, q-bio.NC updates on arXiv.org
-
Thinking in Streaming Video
arXiv:2603.12938v1 Announce Type: cross Abstract: Real-time understanding of continuous video streams is essential for interactive assistants and multimodal agents operating in dynamic environments. However, most existing video reasoning approaches follow a batch paradigm that defers reasoning until the full video context is observed, resulting in high latency and growing computational cost that are incompatible with streaming scenarios. In this paper, we introduce ThinkStream, a framework for
Thinking in Streaming Video
-
cs.AI, q-bio.NC updates on arXiv.org
-
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
arXiv:2603.08262v1 Announce Type: new Abstract: The integration of Large Language Models (LLMs) into the financial domain is driving a paradigm shift from passive information retrieval to dynamic, agentic interaction. While general-purpose tool learning has witnessed a surge in benchmarks, the financial sector, characterized by high stakes, strict compliance, and rapid data volatility, remains critically underserved. Existing financial evaluations predominantly focus on static textual analysis
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
-
cs.AI, q-bio.NC updates on arXiv.org
-
TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable Reward
arXiv:2603.07700v1 Announce Type: cross Abstract: While few-step generative models have enabled powerful image and video generation at significantly lower cost, generic reinforcement learning (RL) paradigms for few-step models remain an unsolved problem. Existing RL approaches for few-step diffusion models strongly rely on back-propagating through differentiable reward models, thereby excluding the majority of important real-world reward signals, e.g., non-differentiable rewards such as humans'
TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable Reward
-
cs.AI, q-bio.NC updates on arXiv.org
-
Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills
arXiv:2512.16301v3 Announce Type: replace Abstract: Large language model (LLM) agents are moving beyond prompting alone. ChatGPT marked the rise of general-purpose LLM assistants, DeepSeek showed that on-policy reinforcement learning with verifiable rewards can improve reasoning and tool use, and OpenClaw highlights a newer direction in which agents accumulate persistent memory and reusable skills. Yet the research landscape remains fragmented across post-training, retrieval, memory, and skill
Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills
-
cs.AI, q-bio.NC updates on arXiv.org
-
Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
arXiv:2603.03485v1 Announce Type: cross Abstract: Recent video diffusion models have achieved impressive capabilities as large-scale generative world models. However, these models often struggle with fine-grained physical consistency, exhibiting physically implausible dynamics over time. In this work, we present \textbf{Phys4D}, a pipeline for learning physics-consistent 4D world representations from video diffusion models. Phys4D adopts \textbf{a three-stage training paradigm} that progressive
Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
-
cs.AI, q-bio.NC updates on arXiv.org
-
GIPO: Gaussian Importance Sampling Policy Optimization
arXiv:2603.03955v1 Announce Type: cross Abstract: Post-training with reinforcement learning (RL) has recently shown strong promise for advancing multimodal agents beyond supervised imitation. However, RL remains limited by poor data efficiency, particularly in settings where interaction data are scarce and quickly become outdated. To address this challenge, GIPO (Gaussian Importance sampling Policy Optimization) is proposed as a policy optimization objective based on truncated importance sampli
GIPO: Gaussian Importance Sampling Policy Optimization
-
cs.AI, q-bio.NC updates on arXiv.org
-
GarmentPile++: Affordance-Driven Cluttered Garments Retrieval with Vision-Language Reasoning
arXiv:2603.04158v1 Announce Type: cross Abstract: Garment manipulation has attracted increasing attention due to its critical role in home-assistant robotics. However, the majority of existing garment manipulation works assume an initial state consisting of only one garment, while piled garments are far more common in real-world settings. To bridge this gap, we propose a novel garment retrieval pipeline that can not only follow language instruction to execute safe and clean retrieval but also g
GarmentPile++: Affordance-Driven Cluttered Garments Retrieval with Vision-Language Reasoning
-
cs.AI, q-bio.NC updates on arXiv.org
-
Generalization of RLVR Using Causal Reasoning as a Testbed
arXiv:2512.20760v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for post-training large language models (LLMs) on complex reasoning tasks. Yet, the conditions under which RLVR yields robust generalization remain underexplored. This paper provides an empirical study of RLVR generalization in the setting of probabilistic inference over causal graphical models. This setting offers two natural axes along which to ex
Generalization of RLVR Using Causal Reasoning as a Testbed
-
cs.AI, q-bio.NC updates on arXiv.org
-
Improving Diffusion Planners by Self-Supervised Action Gating with Energies
arXiv:2603.02650v1 Announce Type: cross Abstract: Diffusion planners are a strong approach for offline reinforcement learning, but they can fail when value-guided selection favours trajectories that score well yet are locally inconsistent with the environment dynamics, resulting in brittle execution. We propose Self-supervised Action Gating with Energies (SAGE), an inference-time re-ranking method that penalises dynamically inconsistent plans using a latent consistency signal. SAGE trains a Joi
Improving Diffusion Planners by Self-Supervised Action Gating with Energies
-
cs.AI, q-bio.NC updates on arXiv.org
-
DMTrack: Spatio-Temporal Multimodal Tracking via Dual-Adapter
arXiv:2508.01592v2 Announce Type: replace-cross Abstract: In this paper, we explore adapter tuning and introduce a novel dual-adapter architecture for spatio-temporal multimodal tracking, dubbed DMTrack. The key of our DMTrack lies in two simple yet effective modules, including a spatio-temporal modality adapter (STMA) and a progressive modality complementary adapter (PMCA) module. The former, applied to each modality alone, aims to adjust spatio-temporal features extracted from a frozen backbo
DMTrack: Spatio-Temporal Multimodal Tracking via Dual-Adapter
-
cs.AI, q-bio.NC updates on arXiv.org
-
VeriStruct: AI-assisted Automated Verification of Data-Structure Modules in Verus
arXiv:2510.25015v4 Announce Type: replace-cross Abstract: We introduce VeriStruct, a novel framework that extends AI-assisted automated verification from single functions to more complex data structure modules in Verus. VeriStruct employs a planner module to orchestrate the systematic generation of abstractions, type invariants, specifications, and proof code. To address the challenge that LLMs often misunderstand Verus' annotation syntax and verification-specific semantics, VeriStruct embeds s
VeriStruct: AI-assisted Automated Verification of Data-Structure Modules in Verus
-
cs.AI, q-bio.NC updates on arXiv.org
-
Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
arXiv:2508.05612v5 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as an effective post-training paradigm for enhancing the reasoning capabilities of multimodal large language model (MLLM). However, current RL pipelines often suffer from training inefficiencies caused by two underexplored issues: Advantage Collapsing, where most advantages in a batch concentrate near zero, and Rollout Silencing, where the proportion of rollouts contributing non-zero gradients dimi
Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
-
cs.AI, q-bio.NC updates on arXiv.org
-
Automated Proof Generation for Rust Code via Self-Evolution
arXiv:2410.15756v3 Announce Type: replace-cross Abstract: Ensuring correctness is crucial for code generation. Formal verification offers a definitive assurance of correctness, but demands substantial human effort in proof construction and hence raises a pressing need for automation. The primary obstacle lies in the severe lack of data-there is much fewer proofs than code snippets for Large Language Models (LLMs) to train upon. In this paper, we introduce SAFE, a framework that overcomes the la
Automated Proof Generation for Rust Code via Self-Evolution
-
cs.AI, q-bio.NC updates on arXiv.org
-
CellINR: Implicitly Overcoming Photo-induced Artifacts in 4D Live Fluorescence Microscopy
arXiv:2508.19300v2 Announce Type: replace-cross Abstract: 4D live fluorescence microscopy is often compromised by prolonged high intensity illumination which induces photobleaching and phototoxic effects that generate photo-induced artifacts and severely impair image continuity and detail recovery. To address this challenge, we propose the CellINR framework, a case-specific optimization approach based on implicit neural representation. The method employs blind convolution and structure amplific