❌

Normal view

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

arXiv:2603.29844v1 Announce Type: cross Abstract: The development of Vision-Language-Action (VLA) models has been significantly accelerated by pre-trained Vision-Language Models (VLMs). However, most existing end-to-end VLAs treat the VLM primarily as a multimodal encoder, directly mapping vision-language features to low-level actions. This paradigm underutilizes the VLM's potential in high-level decision making and introduces training instability, frequently degrading its rich semantic representations. To address these limitations, we introduce DIAL, a framework bridging high-level decision making and low-level motor execution through a differentiable latent intent bottleneck. Specifically, a VLM-based System-2 performs latent world modeling by synthesizing latent visual foresight within the VLM's native feature space; this foresight explicitly encodes intent and serves as the structural bottleneck. A lightweight System-1 policy then decodes this predicted intent together with the current observation into precise robot actions via latent inverse dynamics. To ensure optimization stability, we employ a two-stage training paradigm: a decoupled warmup phase where System-2 learns to predict latent futures while System-1 learns motor control under ground-truth future guidance within a unified feature space, followed by seamless end-to-end joint optimization. This enables action-aware gradients to refine the VLM backbone in a controlled manner, preserving pre-trained knowledge. Extensive experiments on the RoboCasa GR1 Tabletop benchmark show that DIAL establishes a new state-of-the-art, achieving superior performance with 10x fewer demonstrations than prior methods. Furthermore, by leveraging heterogeneous human demonstrations, DIAL learns physically grounded manipulation priors and exhibits robust zero-shot generalization to unseen objects and novel configurations during real-world deployment on a humanoid robot.

Hepatotoxicity Prediction and Multi-omics Reveal Mitochondrial and Lipid Metabolic Dysregulation in PM<sub>2.5</sub>-Induced Liver Fibrosis

26 March 2026 at 18:00

Environ Health (Wash). 2025 Nov 14;4(3):513-521. doi: 10.1021/envhealth.5c00401. eCollection 2026 Mar 20.

ABSTRACT

Prolonged exposure to fine particulate matter (PM2.5) has been linked to chronic liver injury and cancer. However, an alternative risk assessment method to prospective longitudinal studies of exposome-metabolome interactions for liver inflammation-associated hepatocellular carcinoma (HCC) is lacking. This study investigates the risk of long-term real-world PM2.5 exposure in hepatocarcinogenesis through machine learning techniques. Shotgun mass spectrometry (MS) imaging data were acquired from mouse models across a continuum of fibrosis, cirrhosis, and HCC for training a multiclass classification model to identify "No Risk", "Cancer Risk", and "Cancer". Direct infusion-MS data from PM2.5-exposed mouse livers were analyzed to classify risk. By integrating data-driven and knowledge-based approaches, 14 disease progression biomarkers were identified for modeling. Our results suggest that chronic real-world PM2.5 exposure can induce liver fibrosis, presenting cancer risk. Incorporating metabolomics, lipidomics, and transcriptomics, we propose PM2.5 exposure induces mitochondrial dysfunction, activates AMPK signaling, and increases ceramide accumulation, potentially mediating insulin resistance that contributes to nonalcoholic fatty liver disease and HCC progression. This work represents a significant advancement in assessing hepatotoxicity of environmental toxicants by reducing reliance on traditional animal testing methods. It also underscores the potential of emerging technologies in transforming our understanding of PM2.5 exposure, paving the way for targeted interventions.

PMID:41883379 | PMC:PMC13010293 | DOI:10.1021/envhealth.5c00401

❌