❌

Normal view

Show-Harness: Just a VLM Agent Can Play Robots

arXiv:2609.10522v1 Announce Type: cross Abstract: Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VLM directly responsible for fine-grained physical decisions. Through the same interface, Show-Harness demonstrates the feasibility of (1) directly unlocking closed-source frontier VLMs for zero-shot robot control, and (2) adapting small-scale open-source VLMs for low-cost deployment with just a few GPU-hours of fine-tuning. We further develop GUMI (GUI Manipulation Interface), which extends the same semantic action space to GUI-based demonstration collection, allowing humans and agents to "play" robots across embodiments without specialized teleoperation hardware. Extensive experiments show that Show-Harness-equipped VLM agents generalize robustly across tasks, embodiments, and environments, outperforming representative agentic and VLA paradigms. These results suggest that the right interface can unlock substantial embodied capability from foundation VLMs, without requiring additional model capacity or costly embodiment-specific pretraining.

Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression

arXiv:2605.22337v2 Announce Type: replace Abstract: The KV cache used in large language models has linearly growing time complexity, so LLMs face memory blow-up and reduced decoding efficiency when they process long contexts. Current KV Cache eviction has become an important research direction; however, existing methods based on fixed Soft Tokens (e.g., Judge Q) rely on a static parameter set as the query to evaluate the importance of KV pairs, so they cannot adapt dynamically to different input prompts, and they cannot precisely capture complex and changing task relevance. Also, evicted KV pairs are discarded permanently, so this causes irreversible information loss and context breaks. To address this problem, we propose Meta-Soft, a dynamic compression framework based on probe-driven context integration. Specifically, we build a meta-library with a learnable orthogonal basis matrix $\mathcal{L}$, and we use a selector network with Gumbel-Softmax to produce differentiable sparse combination weights, so we dynamically synthesize the most targeted $k$ Soft Tokens from the input prompt features. We append these Soft Tokens to the end of the input sequence to probe key information. We also introduce an attention-flow based integration mechanism, which redistributes the semantic information of removed tokens into retained tokens, and this keeps the dropped context information effectively. Experiments on multiple datasets show that our method outperforms existing state-of-the-art eviction methods and provides a new solution for KV Cache compression.

Integrative multi-omics and experimental validation reveal UBE2C as a central hub gene and prognostic biomarker in hepatocellular carcinoma

Int Immunopharmacol. 2026 May 19;183:116866. doi: 10.1016/j.intimp.2026.116866. Online ahead of print.

ABSTRACT

Hepatocellular carcinoma (HCC) is a lethal malignancy with a high recurrence rate and limited treatment options. Ubiquitin-conjugating enzyme E2 C (UBE2C) is implicated in various cancers, yet its impact on the HCC immune landscape remains incompletely understood. Herein, hub genes in HCC were identified, by integrating co-expression networks and protein-protein interaction analyses, from the TCGA, GEO, and CPTAC databases. Their expression was analysed using a single-cell transcriptomic database and verified in HCC tissues and cell lines via quantitative reverse transcription-PCR and immunoblotting. Functional roles of UBE2C were assessed using in vitro knockdown experiments and an in vivo subcutaneous tumour model. The tumour immune microenvironment was profiled using spatial transcriptomics, RNA-seq data, and ssGSEA. A prognostic nomogram was constructed based on multivariate Cox regression. UBE2C was identified as a significantly upregulated hub gene in HCC. Single-cell RNA-seq revealed predominant expression of UBE2C in hepatocytes, with dynamic upregulation along differentiation trajectories. UBE2C knockdown suppressed proliferation, induced apoptosis, and inhibited tumour growth. Spatial transcriptomics highlighted UBE2C-high regions within proliferative niches exhibiting immunosuppressive traits-including TGFB1 enrichment, impaired CXCL9-CXCR3 signalling, and exclusion of cytotoxic T cells-which were reduced in immunotherapy responders. UBE2C expression correlated with immune checkpoint genes and specific immune cell subsets. A UBE2C-based nomogram integrating T stage and tumour stage robustly predicted patient survival, and miR-300 and miR-381-3p were identified as potential upstream regulators. These findings establish UBE2C as a key driver of HCC progression and a biomarker for prognosis and immunotherapy stratification.

PMID:42155390 | DOI:10.1016/j.intimp.2026.116866

A spatial atlas of the healthy human liver from live donors

Nature, Published online: 15 April 2026; doi:10.1038/s41586-026-10377-y

A human spatial atlas of gene expression in liver based on live donors shows marked porto–central zonation of hepatocytes and non-parenchymal cells, and transcriptomic changes in early steatosis.

RubricBench: Aligning Model-Generated Rubrics with Human Standards

arXiv:2603.01562v2 Announce Type: replace Abstract: As Large Language Model (LLM) alignment evolves from simple completions to complex, highly sophisticated generation, Reward Models are increasingly shifting toward rubric-guided evaluation to mitigate surface-level biases. However, the community lacks a unified benchmark to assess this evaluation paradigm, as existing benchmarks lack both the discriminative complexity and the ground-truth rubric annotations required for rigorous analysis. To bridge this gap, we introduce RubricBench, a curated benchmark with 1,147 pairwise comparisons specifically designed to assess the reliability of rubric-based evaluation. Our construction employs a multi-dimensional filtration pipeline to target hard samples featuring nuanced input complexity and misleading surface bias, augmenting each with expert-annotated, atomic rubrics derived strictly from instructions. Comprehensive experiments reveal a substantial capability gap between human-annotated and model-generated rubrics, indicating that even state-of-the-art models struggle to autonomously specify valid evaluation criteria, lagging considerably behind human-guided performance.

Dataset Distillation via Committee Voting

arXiv:2501.07575v2 Announce Type: replace-cross Abstract: Dataset distillation aims to synthesize a compact yet representative dataset that preserves the essential characteristics of the original data for efficient model training. Existing methods mainly focus on improving data-synthetic alignment or scaling distillation to large datasets. In this work, we propose $\textbf{C}$ommittee $\textbf{V}$oting for $\textbf{D}$ataset $\textbf{D}$istillation ($\textbf{CV-DD}$), an orthogonal approach that leverages the collective knowledge of multiple models to produce higher-quality distilled data. We first establish a strong baseline that achieves state-of-the-art performance through modern architectural and optimization choices. By integrating distributions and predictions from multiple models and generating high-quality soft labels, our method captures a broader range of data characteristics, reduces model-specific bias and the impact of distribution shifts, and significantly improves generalization. This voting-based strategy enhances diversity and robustness, alleviates overfitting, and improves post-evaluation performance. Extensive experiments across multiple datasets and IPC settings demonstrate that CV-DD consistently outperforms single- and multi-model distillation methods and generalizes well to non-training-based frameworks and challenging synthetic-to-real transfer tasks. Code is available at: https://github.com/Jiacheng8/CV-DD.
❌