❌

Normal view

Overcoming missing data in spatial metabolomics with machine learning imputation to accelerate downstream discovery

iScience. 2026 Mar 3;29(4):115203. doi: 10.1016/j.isci.2026.115203. eCollection 2026 Apr 17.

ABSTRACT

Mass spectrometry imaging (MSI)-based spatial metabolomics exhibits extensive missing values; yet, practical guidance on how imputation choices affect both imputation accuracy and downstream spatial analyses remains limited. In this study, we evaluated eight imputation methods, including both existing approaches and a graph convolutional network (GCN)-based method specifically designed for spatial metabolomics data, to identify suitable approaches for spatial metabolomics. To enable comprehensive assessment, we developed an evaluation framework focusing on two objective criteria: (a) imputation accuracy and (b) preservation of spatial cluster structure. We assembled six benchmark datasets spanning mouse brain and liver, human kidney and stomach, and plant seed sections, and conducted controlled dropout simulations of missing values. Across both evaluation dimensions, including imputation accuracy and preservation of spatial cluster structure, RF ranked first overall, and GCN ranked second in both dimensions. Overall, this systematic, dual-perspective benchmark study provides guidance for selecting imputation strategies in spatial metabolomics research.

PMID:41869568 | PMC:PMC12999350 | DOI:10.1016/j.isci.2026.115203

AdaCultureSafe: Adaptive Cultural Safety Grounded by Cultural Knowledge in Large Language Models

arXiv:2603.08275v1 Announce Type: cross Abstract: With the widespread adoption of Large Language Models (LLMs), respecting indigenous cultures becomes essential for models' culturally safety and responsible global applications. Existing studies separately consider cultural safety and cultural knowledge and neglect that the former should be grounded by the latter. This severely prevents LLMs from yielding culture-specific respectful responses. Consequently, adaptive cultural safety remains a formidable task. In this work, we propose to jointly model cultural safety and knowledge. First and foremost, cultural-safety and knowledge-paired data serve as the key prerequisite to conduct this research. However, the cultural diversity across regions and the subtlety of cultural differences pose significant challenges to the creation of such paired evaluation data. To address this issue, we propose a novel framework that integrates authoritative cultural knowledge descriptions curation, LLM-automated query generation, and heavy manual verification. Accordingly, we obtain a dataset named AdaCultureSafe containing 4.8K manually decomposed fine-grained cultural descriptions and the corresponding 48K manually verified safety- and knowledge-oriented queries. Upon the constructed dataset, we evaluate three families of popular LLMs on their cultural safety and knowledge proficiency, via which we make a critical discovery: no significant correlation exists between their cultural safety and knowledge proficiency. We then delve into the utility-related neuron activations within LLMs to investigate the potential cause of the absence of correlation, which can be attributed to the difference of the objectives of pre-training and post-alignment. We finally present a knowledge-grounded method, which significantly enhances cultural safety by enforcing the integration of knowledge into the LLM response generation process.
❌