❌

Normal view

GenOM: Ontology Matching with Description Generation and Large Language Model

arXiv:2508.10703v3 Announce Type: replace Abstract: Ontology matching (OM) plays an essential role in enabling semantic interoperability and integration across heterogeneous knowledge sources, particularly in the biomedical domain which contains numerous complex concepts related to diseases and pharmaceuticals. This paper introduces GenOM, a large language model (LLM)-based ontology alignment framework, which enriches the semantic representations of ontology concepts via generating textual definitions, retrieves alignment candidates with an embedding model, and incorporates exact matching-based tools to improve precision. Extensive experiments conducted on the OAEI Bio-ML track demonstrate that GenOM can often achieve competitive performance, surpassing many baselines including traditional OM systems and recent LLM-based methods. Further ablation studies confirm the effectiveness of semantic enrichment and few-shot prompting, highlighting the framework's robustness and adaptability.

Dynamic Targetable Extracellular Vesicle Surface Proteins Monitor Depth of Response to CAR T Therapy

Res Sq [Preprint]. 2026 Mar 18:rs.3.rs-8913641. doi: 10.21203/rs.3.rs-8913641/v1.

ABSTRACT

Extracellular vesicles (EVs) represent a promising liquid biopsy platform in multiple myeloma (MM). We developed an MM EV Surface Protein Assay to quantify and dynamically monitor four MM EV subpopulations defined by targetable MM surface proteins (BCMA, CD38, GPRC5D, and CD319) across 336 serial blood samples from 45 relapsed/refractory MM (RRMM) patients treated with anti-BCMA chimeric antigen receptor (CAR) T-cell therapy. All four MM EV subpopulations significantly decreased in 43 patients with initial response, while BCMA+, GPRC5D+, and CD319+ MM EVs increased in 19 patients with progression, and antigen escape was detected by BCMA+ MM EVs. MM EV subpopulations differentiated minimal residual disease (MRD) status and complemented MRD for detecting early relapse before clinical progression. Notably, CD319+ MM EVs were early predictors of progression-free and overall survival in MRD-negative patients. This assay enables noninvasive monitoring of deep response, progression, and antigen escape, and stratifies survival in MRD-negative patients with RRMM.

PMID:41890853 | PMC:PMC13015583 | DOI:10.21203/rs.3.rs-8913641/v1

Profiling of the mycobiome and metabolome: a comparative study of benign pulmonary nodules and lung adenocarcinoma

Front Cell Infect Microbiol. 2026 Feb 23;16:1732958. doi: 10.3389/fcimb.2026.1732958. eCollection 2026.

ABSTRACT

INTRODUCTION: Lung adenocarcinoma (LUAD), the most common subtype of non-small cell lung cancer, is a form of malignant pulmonary nodule that requires clinical differentiation from benign pulmonary nodules (BPN). The mechanisms underlying the development of LUAD are complex, and effective non-invasive methods for differentiating BPN from LUAD are lacking. This study aimed not only to distinguish BPN from LUAD using gut fungi and serum metabolites, but also to establish an integrated network of gut fungi-metabolite-cytokine interactions.

METHODS: Fecal and serum samples from individuals with BPN and patients with LUAD were subjected to internal transcribed spacer sequencing, ultra-performance liquid chromatography-tandem mass spectrometry, and multiplex Luminex assays to quantify gut fungi, metabolites, and cytokines, respectively.

RESULTS: A significant difference in gut fungal communities was observed between the BPN and LUAD groups. Multiple genera and species were more abundant in LUAD than in BPN. Docosapentaenoic acid n-6 (DPAn-6), indole-3-propionic acid (IPA), and interferon-γ-induced protein 10 (IP-10) were significantly elevated in the LUAD group. The integrated model established using a combination of gut fungi and metabolites demonstrated excellent performance in distinguishing BPN from LUAD. A network of interactions was established among differentially abundant gut fungi, serum metabolites, and cytokines.

CONCLUSION: Our study identifies a novel panel of fungal and metabolite biomarkers for differentiating between BPN and LUAD, and constructs a multi-omics network that provides new insights into investigating the mechanistic role of gut mycobiota dysbiosis in LUAD.

PMID:41809995 | PMC:PMC12968269 | DOI:10.3389/fcimb.2026.1732958

Local Shapley: Model-Induced Locality and Optimal Reuse in Data Valuation

arXiv:2603.03672v1 Announce Type: cross Abstract: The Shapley value provides a principled foundation for data valuation, but exact computation is #P-hard due to the exponential coalition space. Existing accelerations remain global and ignore a structural property of modern predictors: for a given test instance, only a small subset of training points influences the prediction. We formalize this model-induced locality through support sets defined by the model's computational pathway (e.g., neighbors in KNN, leaves in trees, receptive fields in GNNs), showing that Shapley computation can be projected onto these supports without loss when locality is exact. This reframes Shapley evaluation as a structured data processing problem over overlapping support-induced subset families rather than exhaustive coalition enumeration. We prove that the intrinsic complexity of Local Shapley is governed by the number of distinct influential subsets, establishing an information-theoretic lower bound on retraining operations. Guided by this result, we propose LSMR (Local Shapley via Model Reuse), an optimal subset-centric algorithm that trains each influential subset exactly once via support mapping and pivot scheduling. For larger supports, we develop LSMR-A, a reuse-aware Monte Carlo estimator that remains unbiased with exponential concentration, with runtime determined by the number of distinct sampled subsets rather than total draws. Experiments across multiple model families demonstrate substantial retraining reductions and speedups while preserving high valuation fidelity.

Unifying Evolutionary Prompt Search and Reinforcement Learning for LLM Self-Improvement

arXiv:2602.14697v2 Announce Type: replace Abstract: Building agentic systems that can autonomously self-improve from experience is a longstanding goal of AI. Large language models (LLMs) today primarily self-improve via two mechanisms: self-reflection for context updates, and reinforcement learning (RL) for weight updates. In this work, we propose Evolutionary System Prompt Learning (E-SPL), a method for jointly improving model contexts and model weights. In each RL iteration, E-SPL samples trajectories under multiple system prompts in parallel. It applies RL updates to LLM weights conditioned on system prompts, and evolutionary updates to system prompts via mutation and crossover, two genetic operators based on LLM self-reflection. Each system prompt is assigned a TrueSkill rating for evolutionary selection, updated from relative performance within each RL iteration. E-SPL encourages a natural division between declarative knowledge encoded in prompts and procedural knowledge encoded in weights, resulting in improved performance across reasoning and agentic tasks. For instance, in an easy-to-hard (AIME $\rightarrow$ BeyondAIME) generalization setting, E-SPL improves RL success rate from 38.8% $\rightarrow$ 45.1% while also outperforming reflective prompt evolution (40.0%). Overall, our results demonstrate that RL and evolutionary prompt search are deeply synergistic, and unifying the two yields consistent gains in sample efficiency and generalization. Code: https://github.com/LunjunZhang/E-SPL
❌