❌

Reading view

Online Video Agent Harness for Long Video Understanding

arXiv:2609.12818v1 Announce Type: cross Abstract: Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal spans, while packing dense frames into a single VLM context incurs \textit{context rot} and high cost. Existing video agents often rely on query-agnostic offline preprocessing or ad hoc tool sets, which can miss query-specific details and waste computation. In this work, we present VideoXAgent, a purely online video-agent harness for long video understanding that starts from the given video file and user query, plans and decomposes the task, invokes specialized expert tools on demand, and aggregates multimodal evidence to produce a final answer while resolving conflicts among observations. To support this on-demand invocation, we design a suite of heterogeneous expert tools guided by a data-driven taxonomy of atomic capabilities, spanning scripts, VLMs, and domain models (e.g., detection, OCR, ASR, face recognition). The harness further enforces objective evidence prompting and budget-aware control to curb hallucination and non-termination. Across Video-MME-Long, LongVideoBench-Long, LVBench, and MINERVA, VideoXAgent is competitive with frontier LMMs and video agents under a smaller context footprint---about 50k tokens of agent context per sample, even on hour-long videos. In particular, on complex video-reasoning benchmarks such as MINERVA, it matches this level while using only about 15\% of the context of a 1,024-frame dense-packing baseline. Notably, the harness remains effective with a visually weak or even text-only orchestrator, suggesting that strong long-video understanding can emerge from progressive agentic evidence seeking rather than from packing the full video into a single context. Project page: https://go-agent-x.github.io/video_agent_harness/
  •  

Overcoming missing data in spatial metabolomics with machine learning imputation to accelerate downstream discovery

iScience. 2026 Mar 3;29(4):115203. doi: 10.1016/j.isci.2026.115203. eCollection 2026 Apr 17.

ABSTRACT

Mass spectrometry imaging (MSI)-based spatial metabolomics exhibits extensive missing values; yet, practical guidance on how imputation choices affect both imputation accuracy and downstream spatial analyses remains limited. In this study, we evaluated eight imputation methods, including both existing approaches and a graph convolutional network (GCN)-based method specifically designed for spatial metabolomics data, to identify suitable approaches for spatial metabolomics. To enable comprehensive assessment, we developed an evaluation framework focusing on two objective criteria: (a) imputation accuracy and (b) preservation of spatial cluster structure. We assembled six benchmark datasets spanning mouse brain and liver, human kidney and stomach, and plant seed sections, and conducted controlled dropout simulations of missing values. Across both evaluation dimensions, including imputation accuracy and preservation of spatial cluster structure, RF ranked first overall, and GCN ranked second in both dimensions. Overall, this systematic, dual-perspective benchmark study provides guidance for selecting imputation strategies in spatial metabolomics research.

PMID:41869568 | PMC:PMC12999350 | DOI:10.1016/j.isci.2026.115203

  •  
❌