❌

Normal view

Explainable Molecular Property Prediction: Aligning Chemical Concepts with Predictions via Language Models

arXiv:2405.16041v4 Announce Type: replace-cross Abstract: Providing explainable molecular property predictions is critical for many scientific domains, such as drug discovery and material science. Though transformer-based language models have shown great potential in accurate molecular property prediction, they neither provide chemically meaningful explanations nor faithfully reveal the molecular structure-property relationships. In this work, we develop a framework for explainable molecular property prediction based on language models, dubbed as Lamole, which can provide chemical concepts-aligned explanations. We take a string-based molecular representation -- Group SELFIES -- as input tokens to pretrain and fine-tune our Lamole, as it provides chemically meaningful semantics. By disentangling the information flows of Lamole, we propose combining self-attention weights and gradients for better quantification of each chemically meaningful substructure's impact on the model's output. To make the explanations more faithfully respect the structure-property relationship, we then carefully craft a marginal loss to explicitly optimize the explanations to be able to align with the chemists' annotations. We bridge the manifold hypothesis with the elaborated marginal loss to prove that the loss can align the explanations with the tangent space of the data manifold, leading to concept-aligned explanations. Experimental results over six mutagenicity datasets and one hepatotoxicity dataset demonstrate Lamole can achieve comparable classification accuracy and boost the explanation accuracy by up to 14.3%, being the state-of-the-art in explainable molecular property prediction.

Targeted inhibition of gastric adenocarcinoma by nano-curcumin liposomes: Insights from combined machine learning and experimental analyses into the mechanisms of cuproptosis and metabolic reprogramming

Int J Pharm. 2025 Nov 9:126368. doi: 10.1016/j.ijpharm.2025.126368. Online ahead of print.

ABSTRACT

PURPOSE: Gastric adenocarcinoma is a highly aggressive malignancy characterized by a complex tumor microenvironment. Nano-curcumin liposomes hold great potential in inhibiting tumor growth and survival, as well as inducing cuproptosis and oxidative stress. Although the anticancer properties of curcumin have been demonstrated, the specific mechanisms by which curcumin inhibites gastric adenocarcinoma through cuproptosis remains unclear. This study investigated how nano-curcumin liposomes mediated the inhibition of gastric adenocarcinoma cell proliferation and survival via cuproptosis.

METHODS: This study utilized the gastric adenocarcinoma cell line AGS to establish 2D and 3D in vitro gastric adenocarcinoma models. Furthermore, we prepared nano-curcumin liposomes to investigate their effects and regulatory mechanisms on AGS gastric adenocarcinoma models. A series of in vitro assays, including flow cytometry, CCK-8, scratch assays and morphological assessments, were performed to evaluate the effects of nano-curcumin liposomes on cell apoptosis, proliferation and migration. Additionally, bioinformatics and machine learning methods were employed to identify key targets that inhibited gastric adenocarcinoma growth and survival associated with nano-curcumin liposomes, which were further validated through RT-qPCR and omics analysis. Computer simulations were also conducted to assess the stability of binding interactions between curcumin and key target proteins.

RESULTS: Cellular experiments demonstrated that nano-curcumin liposomes significantly inhibited proliferation and invasive capacity of gastric adenocarcinoma cells while promoting cellular oxidative stress. Bioinformatics and machine learning analyses identified FDX1, GPX4, SERPINE1 and SLC27A5 as key targets. RT-qPCR results confirmed that nano-curcumin liposomes significantly downregulated the expression of these targets. Molecular dynamics simulations indicated that curcumin could form stable binding interactions with key protein targets.

CONCLUSION: This study revealed that nano-curcumin liposomes inhibited growth and survival of gastric adenocarcinoma cells by interfering with the expression of FDX1, GPX4, SERPINE1 and SLC27A5, which were closely linked to copper-induced oxidative stress. Nano-curcumin liposomes downregulated the expression of FDX1 and GPX4, disrupted mitochondrial energy metabolism, and induced oxidative stress, thereby promoting tumor-associated programmed cell death linked to cuproptosis. Furthermore, by downregulating SERPINE1, nano-curcumin liposomes modulated cell adhesion and migration, inhibiting the invasive and metastatic potential of tumor cells. Finally, downregulation of SLC27A5 altered tumor metabolism and cellular homeostasis, induced oxidative stress, and disrupted intracellular environmental stability, thereby suppressing the growth of gastric adenocarcinoma.

PMID:41218732 | DOI:10.1016/j.ijpharm.2025.126368

HiF-DTA: Hierarchical Feature Learning Network for Drug-Target Affinity Prediction

arXiv:2510.27281v1 Announce Type: cross Abstract: Accurate prediction of Drug-Target Affinity (DTA) is crucial for reducing experimental costs and accelerating early screening in computational drug discovery. While sequence-based deep learning methods avoid reliance on costly 3D structures, they still overlook simultaneous modeling of global sequence semantic features and local topological structural features within drugs and proteins, and represent drugs as flat sequences without atomic-level, substructural-level, and molecular-level multi-scale features. We propose HiF-DTA, a hierarchical network that adopts a dual-pathway strategy to extract both global sequence semantic and local topological features from drug and protein sequences, and models drugs multi-scale to learn atomic, substructural, and molecular representations fused via a multi-scale bilinear attention module. Experiments on Davis, KIBA, and Metz datasets show HiF-DTA outperforms state-of-the-art baselines, with ablations confirming the importance of global-local extraction and multi-scale fusion.

A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications

arXiv:2510.16724v2 Announce Type: replace Abstract: The advent of large language models (LLMs) has transformed information access and reasoning through open-ended natural language interaction. However, LLMs remain limited by static knowledge, factual hallucinations, and the inability to retrieve real-time or domain-specific information. Retrieval-Augmented Generation (RAG) mitigates these issues by grounding model outputs in external evidence, but traditional RAG pipelines are often single turn and heuristic, lacking adaptive control over retrieval and reasoning. Recent advances in agentic search address these limitations by enabling LLMs to plan, retrieve, and reflect through multi-step interaction with search environments. Within this paradigm, reinforcement learning (RL) offers a powerful mechanism for adaptive and self-improving search behavior. This survey provides the first comprehensive overview of \emph{RL-based agentic search}, organizing the emerging field along three complementary dimensions: (i) What RL is for (functional roles), (ii) How RL is used (optimization strategies), and (iii) Where RL is applied (scope of optimization). We summarize representative methods, evaluation protocols, and applications, and discuss open challenges and future directions toward building reliable and scalable RL driven agentic search systems. We hope this survey will inspire future research on the integration of RL and agentic search. Our repository is available at https://github.com/ventr1c/Awesome-RL-based-Agentic-Search-Papers.

Decoding the tumor immune microenvironment in lung squamous cell carcinoma: characteristics, regulatory mechanisms, and future directions in immunotherapy

24 October 2025 at 18:00

Transl Lung Cancer Res. 2025 Sep 30;14(9):4112-4130. doi: 10.21037/tlcr-2025-350. Epub 2025 Sep 18.

ABSTRACT

Lung squamous cell carcinoma (LUSC), a predominant type of lung cancer, is marked by an unfavorable prognosis and limited therapeutic options. Unlike lung adenocarcinoma (LUAD), LUSC exhibits few driver mutations, resulting in minimal benefits from targeted therapies for these patients. Despite the transformative effects of immunotherapy on patient outcomes, only a subset of patients achieving durable responses. This heterogeneity in treatment outcomes is increasingly attributed to the complex feature of the tumor immune microenvironment (TIME) in LUSC. The TIME of LUSC is a highly dynamic ecosystem composed of diverse immune cell populations and stromal components that collectively foster an immune-evasive niche. Recent breakthroughs in multi-omics technologies, particularly single-cell RNA sequencing (scRNA-seq) and spatial omics, have provided unprecedented resolution in dissecting the cellular and molecular architecture of the TIME in LUSC. These technologies have enabled the identification of distinct immune cells and their spatial interactions with the tumor, shedding light on the mechanisms underlying immune evasion and resistance to immunotherapy. Building on these advancements, this review establishes a new classification of the TIME which may guide patient stratification and personalized immunotherapy. And we comprehensively offer a detailed examination of the principal characteristics and regulatory mechanisms of the TIME, highlighting potential immunotherapeutic strategies tailored to this distinct immunological context.

PMID:41133013 | PMC:PMC12541881 | DOI:10.21037/tlcr-2025-350

A multiomics dataset of paired CT image and plasma cell-free DNA end motif for patients with pulmonary nodules

Sci Data. 2025 Apr 1;12(1):545. doi: 10.1038/s41597-025-04912-1.

ABSTRACT

Diagnosing lung cancer at a curable stage offers the opportunity for a favorable prognosis. The emerging epigenomics analysis on plasma cell-free DNA (cfDNA), including 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) modifications, has acted as a promising approach facilitating the identification of lung cancer. And, integrating 5mC biomarker with chest computed tomography (CT) image features could optimize the diagnosis of lung cancer, exceeding the performance of models built on single feature. However, the clinical applicability of integrated markers might be limited by the potential risk of overfitting due to small sample size. Hence, we prospectively collected peripheral blood sample and the paired chest CT images of 2032 patients with indeterminate pulmonary nodules across 5 centers, and constructed a large-scale, multi-institutional, multiomics database that encompass CT imaging data and plasma cfDNA fragmentomic in 5mC-, 5hmC-enriched regions. To our best knowledge, this dataset is the first radio-epigenomic dataset with the largest sample size, and provides multi-dimensional insights for early diagnosis of lung cancer, facilitating the individuated management for lung cancer.

PMID:40169596 | PMC:PMC11961589 | DOI:10.1038/s41597-025-04912-1

WNT2–SOX4 positive feedback loop promotes chemoresistance and tumorigenesis by inducing stem-cell like properties in gastric cancer

Oncogene, Published online: 26 August 2023; doi:10.1038/s41388-023-02816-1

WNT2–SOX4 positive feedback loop promotes chemoresistance and tumorigenesis by inducing stem-cell like properties in gastric cancer

Methylation of HBP1 by PRMT1 promotes tumor progression by regulating actin cytoskeleton remodeling

Oncogenesis, Published online: 08 August 2022; doi:10.1038/s41389-022-00421-7

Methylation of HBP1 by PRMT1 promotes tumor progression by regulating actin cytoskeleton remodeling
❌