❌

Normal view

3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding

arXiv:2603.23447v1 Announce Type: cross Abstract: While multi-modality large language models excel in object-centric or indoor scenarios, scaling them to 3D city-scale environments remains a formidable challenge. To bridge this gap, we propose 3DCity-LLM, a unified framework designed for 3D city-scale vision-language perception and understanding. 3DCity-LLM employs a coarse-to-fine feature encoding strategy comprising three parallel branches for target object, inter-object relationship, and global scene. To facilitate large-scale training, we introduce 3DCity-LLM-1.2M dataset that comprises approximately 1.2 million high-quality samples across seven representative task categories, ranging from fine-grained object analysis to multi-faceted scene planning. This strictly quality-controlled dataset integrates explicit 3D numerical information and diverse user-oriented simulations, enriching the question-answering diversity and realism of urban scenarios. Furthermore, we apply a multi-dimensional protocol based on text-similarity metrics and LLM-based semantic assessment to ensure faithful and comprehensive evaluations for all methods. Extensive experiments on two benchmarks demonstrate that 3DCity-LLM significantly outperforms existing state-of-the-art methods, offering a promising and meaningful direction for advancing spatial reasoning and urban intelligence. The source code and dataset are available at https://github.com/SYSU-3DSTAILab/3D-City-LLM.

SIRT3 deacetylates STEAP4 to modulate cuproptosis sensitivity via mitochondrial metabolic reprogramming in HBV-related HCC

Cell Death Differ. 2026 Mar 16. doi: 10.1038/s41418-026-01713-w. Online ahead of print.

ABSTRACT

Hepatitis B virus (HBV) infection remains a leading etiological driver of hepatocellular carcinoma (HCC). Cuproptosis is a recently defined copper-dependent form of regulated cell death that selectively eliminates mitochondria-dependent cells; whether HBV rewires this vulnerability remains unknown. Here we unveil a novel HBV X protein (HBx)-driven mechanism of cuproptosis evasion. Integrative analysis of clinical specimens, HBx-transgenic (HBx-Tg) mice, and multi-omics datasets revealed marked downregulation of STEAP4 (six-transmembrane epithelial antigen of prostate 4), a metalloreductase essential for cuproptosis sensitivity, in HBV-positive HCC. Mechanistically, HBx attenuates sirtuin 3 (SIRT3), impairing deacetylation of STEAP4 at lysine 404 and abolishing its mitochondrial targeting. Consequently, cells switch from the tricarboxylic acid (TCA) cycle respiration to glycolysis, reducing sensitivity to the copper ionophore elesclomol (ES). Restoring STEAP4 expression or pharmacological activation of SIRT3 with honokiol (HKL) re-instated mitochondrial STEAP4 localization and re-sensitized HBV-related HCC cells to cuproptosis; combination with ES produced synergistic tumor suppression in vitro and in orthotopic models. Collectively, our findings establish the SIRT3-STEAP4 axis as a novel regulator of cuproptosis resistance in HBV-related HCC. HBx-mediated repression of SIRT3 disrupts STEAP4 deacetylation and mitochondrial targeting, fostering metabolic reprogramming and evasion of copper-induced cell death. The results provide a pre-clinical rationale for copper-directed combination strategies in HBV-associated HCC.

PMID:41840161 | DOI:10.1038/s41418-026-01713-w

PROSPECT: Unified Streaming Vision-Language Navigation via Semantic--Spatial Fusion and Latent Predictive Representation

arXiv:2603.03739v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have advanced zero-shot end-to-end Vision-Language Navigation (VLN), yet robust navigation requires not only semantic understanding but also predictive modeling of environment dynamics and spatial structure. We propose PROSPECT, a unified streaming navigation agent that couples a streaming Vision-Language-Action (VLA) policy with latent predictive representation learning. PROSPECT uses CUT3R as a streaming 3D foundation spatial encoder to produce long-context, absolute-scale spatial features, and fuses them with SigLIP semantic features via cross-attention. During training, we introduce learnable stream query tokens that query the streaming context and predict next-step 2D and 3D latent features (rather than pixels or explicit modalities), supervised in the latent spaces of frozen SigLIP and CUT3R teachers. The predictive branch shapes internal representations without inference overhead. Experiments on VLN-CE benchmarks and real-robot deployment demonstrate state-of-the-art performance and improved long-horizon robustness under diverse lighting. We will release code for the community soon.

Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy

arXiv:2507.01352v3 Announce Type: replace-cross Abstract: Despite the critical role of reward models (RMs) in Reinforcement Learning from Human Feedback (RLHF), current state-of-the-art open RMs perform poorly on most existing evaluation benchmarks, failing to capture nuanced human preferences. We hypothesize that this brittleness stems primarily from limitations in preference datasets, which are often narrowly scoped, synthetically labeled, or lack rigorous quality control. To address these challenges, we present SynPref-40M, a large-scale preference dataset comprising 40 million preference pairs. To enable data curation at scale, we design a human-AI synergistic two-stage pipeline that leverages the complementary strengths of human annotation quality and AI scalability. In this pipeline, humans provide verified annotations, while LLMs perform automatic curation based on human guidance. Training on this preference mixture, we introduce Skywork-Reward-V2, a suite of eight reward models ranging from 0.6B to 8B parameters, trained on a carefully curated subset of 26 million preference pairs from SynPref-40M. We demonstrate that Skywork-Reward-V2 is versatile across a wide range of capabilities, including alignment with human preferences, objective correctness, safety, resistance to stylistic biases, and best-of-N scaling. These reward models achieve state-of-the-art performance across seven major reward model benchmarks, outperform generative reward models, and demonstrate strong downstream performance. Ablation studies confirm that effectiveness stems not only from data scale but also from high-quality curation. The Skywork-Reward-V2 series represents substantial progress in open reward models, demonstrating how human-AI curation synergy can unlock significantly higher data quality.
❌