❌

Reading view

SHOE: Semantic HOI Open-Vocabulary Evaluation Metric

arXiv:2604.01586v1 Announce Type: cross Abstract: Open-vocabulary human-object interaction (HOI) detection is a step towards building scalable systems that generalize to unseen interactions in real-world scenarios and support grounded multimodal systems that reason about human-object relationships. However, standard evaluation metrics, such as mean Average Precision (mAP), treat HOI classes as discrete categorical labels and fail to credit semantically valid but lexically different predictions (e.g., "lean on couch" vs. "sit on couch"), limiting their applicability for evaluating open-vocabulary predictions that go beyond any predefined set of HOI labels. We introduce SHOE (Semantic HOI Open-Vocabulary Evaluation), a new evaluation framework that incorporates semantic similarity between predicted and ground-truth HOI labels. SHOE decomposes each HOI prediction into its verb and object components, estimates their semantic similarity using the average of multiple large language models (LLMs), and combines them into a similarity score to evaluate alignment beyond exact string match. This enables a flexible and scalable evaluation of both existing HOI detection methods and open-ended generative models using standard benchmarks such as HICO-DET. Experimental results show that SHOE scores align more closely with human judgments than existing metrics, including LLM-based and embedding-based baselines, achieving an agreement of 85.73% with the average human ratings. Our work underscores the need for semantically grounded HOI evaluation that better mirrors human understanding of interactions. We will release our evaluation metric to the public to facilitate future research.
  •  

Lifting Unlabeled Internet-level Data for 3D Scene Understanding

arXiv:2604.01907v1 Announce Type: cross Abstract: Annotated 3D scene data is scarce and expensive to acquire, while abundant unlabeled videos are readily available on the internet. In this paper, we demonstrate that carefully designed data engines can leverage web-curated, unlabeled videos to automatically generate training data, to facilitate end-to-end models in 3D scene understanding alongside human-annotated datasets. We identify and analyze bottlenecks in automated data generation, revealing critical factors that determine the efficiency and effectiveness of learning from unlabeled data. To validate our approach across different perception granularities, we evaluate on three tasks spanning low-level perception, i.e., 3D object detection and instance segmentation, to high-evel reasoning, i.e., 3D spatial Visual Question Answering (VQA) and Vision-Lanugage Navigation (VLN). Models trained on our generated data demonstrate strong zero-shot performance and show further improvement after finetuning. This demonstrates the viability of leveraging readily available web data as a path toward more capable scene understanding systems.
  •  

Efficient Reasoning with Balanced Thinking

arXiv:2603.12372v3 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they often suffer from overthinking, expending redundant computational steps on simple problems, or underthinking, failing to explore sufficient reasoning paths despite inherent capabilities. These issues lead to inefficiencies and potential inaccuracies, limiting practical deployment in resource-constrained settings. Existing methods to mitigate overthinking, such as suppressing reflective keywords or adjusting reasoning length, may inadvertently induce underthinking, compromising accuracy. Therefore, we propose ReBalance, a training-free framework that achieves efficient reasoning with balanced thinking. ReBalance leverages confidence as a continuous indicator of reasoning dynamics, identifying overthinking through high confidence variance and underthinking via consistent overconfidence. By aggregating hidden states from a small-scale dataset into reasoning mode prototypes, we compute a steering vector to guide LRMs' reasoning trajectories. A dynamic control function modulates this vector's strength and direction based on real-time confidence, pruning redundancy during overthinking, and promoting exploration during underthinking. Extensive experiments conducted on four models ranging from 0.5B to 32B, and across nine benchmarks in math reasoning, general question answering, and coding tasks demonstrate that ReBalance effectively reduces output redundancy while improving accuracy, offering a general, training-free, and plug-and-play strategy for efficient and robust LRM deployment. Project page and code are available at https://rebalance-ai.github.io .
  •  

Distinctive respiratory toxicity induced by hypoxanthine metabolic disorder from polystyrene microplastics and nanoplastics at environmentally relevant doses: multi-omics insights and experimental validation

Environ Int. 2026 Mar 28;210:110212. doi: 10.1016/j.envint.2026.110212. Online ahead of print.

ABSTRACT

Microplastics (MPs) and nanoplastics (NPs) are pervasive environmental contaminants, raising concerns about their potential to cause inflammation, oxidative stress, and lung injury through respiratory toxicity. Due to their smaller size, larger surface area, and greater reactivity, NPs may pose a greater risk than MPs, yet size-dependent toxicity mechanisms remain unclear. This study investigates the distinct early molecular initiating events and toxicological effects of 1 ΞΌm polystyrene MPs (PS-MPs) and 20 nm polystyrene NPs (PS-NPs). Based on the internal exposure dose estimated from Py-GC/MS analysis, in vitro exposure concentrations were set at 0, 62.5, 125, 250, 500, and 1000 ΞΌg/mL. Multi-omics sequencing and integrative analysis identify specific proteomic and metabolomic alterations. Molecular dynamics simulations and co-immunoprecipitation assays elucidate binding interactions between PS-NPs-induced proteins and metabolic enzymes. In vitro and in vivo experiments reveal a greater accumulation of PS-NPs through endocytosis compared to PS-MPs; while pronounced histopathological damage with inflammatory response in mice lungs were only induced by PS-NPs, rather than PS-MPs. Compared to control group, PS-MPs partly caused proteomic or metabolomic perturbations, while PS-NPs induced significant differential expression of more extensive proteins and metabolites. PS-NPs exposure specifically upregulates insulin-like growth factor 2 receptor (IGF2R) expression and reduces Hypoxanthine levels when compared with PS-MPs. IGF2R directly interacts with Hypoxanthine-guanine phosphoribosyl transferase (HPRT), a key enzyme in Hypoxanthine metabolism, causing its disruption. This study provides important insights into the comparative toxic effects between PS-NPs with PS-MPs, especially the unique toxicological mechanisms of PS-NPs, thereby advancing the understanding of airborne plastic pollutant risks and supporting future regulatory assessments.

PMID:41921402 | DOI:10.1016/j.envint.2026.110212

  •  
❌