❌

Reading view

LICO: Large Language Models for In-Context Molecular Optimization

arXiv:2406.18851v2 Announce Type: replace-cross Abstract: Optimizing black-box functions is a fundamental problem in science and engineering. To solve this problem, many approaches learn a surrogate function that estimates the underlying objective from limited historical evaluations. Large Language Models (LLMs), with their strong pattern-matching capabilities via pretraining on vast amounts of data, stand out as a potential candidate for surrogate modeling. However, directly prompting a pretrained language model to produce predictions is not feasible in many scientific domains due to the scarcity of domain-specific data in the pretraining corpora and the challenges of articulating complex problems in natural language. In this work, we introduce LICO, a general-purpose model that extends arbitrary base LLMs for black-box optimization, with a particular application to the molecular domain. To achieve this, we equip the language model with a separate embedding layer and prediction layer, and train the model to perform in-context predictions on a diverse set of functions defined over the domain. Once trained, LICO can generalize to unseen molecule properties simply via in-context prompting. LICO performs competitively on PMO, a challenging molecular optimization benchmark comprising 23 objective functions, and achieves state-of-the-art performance on its low-budget version PMO-1K.
  •  

MLG2Net: Molecular Global Graph Network for Drug Response Prediction in Lung Cancer Cell Lines

J Med Syst. 2025 Apr 10;49(1):47. doi: 10.1007/s10916-025-02182-3.

ABSTRACT

Drug response prediction (DRP) is a central task in the era of precision medicine. Over the past decade, the emergence of deep learning (DL) has greatly contributed to addressing DRP challenges. Notably, the prediction of DRP for cancer cell lines benefits significantly from data availability for model development. However, an effective predictive model is still challenging due to issues with data quality, high-dimensional data, and multi-omics data integration. In this study, we introduce MLG2Net, a deep-learning model inspired by graph neural networks designed to predict DRP in lung cancer cell lines based on pharmacogenomics data. Our model comprises two key components: drug SMILES described by local and global graph networks and cell line genomics are illustrated as a map. Our results show that MLG2Net outperforms three reference graph networks. MLG2Net performance reached a Pearson coefficient correlation ( C C p ) of 0.8616 and a root mean square error (RMSE) of 2.94e-6 in predicting drug responses for Lung Adenocarcinoma (LUAD) cell lines. Subsequent testing on the Lung Squamous Cell Carcinoma (LUSC) dataset reveals lower performance ( C C p : 0.7999, RMSE: 4.08e-6), attributed to the dataset's smaller size influencing model capacity. Moreover, we assessed the model's architecture by isolating its components, with results indicating that the global network is particularly effective in this task. In conclusion, MLG2Net exhibited promising applications in DRP for cancer cell lines, with potential advancements by incorporating larger datasets.

PMID:40208442 | DOI:10.1007/s10916-025-02182-3

  •  

MLG2Net: Molecular Global Graph Network for Drug Response Prediction in Lung Cancer Cell Lines

J Med Syst. 2025 Apr 10;49(1):47. doi: 10.1007/s10916-025-02182-3.

ABSTRACT

Drug response prediction (DRP) is a central task in the era of precision medicine. Over the past decade, the emergence of deep learning (DL) has greatly contributed to addressing DRP challenges. Notably, the prediction of DRP for cancer cell lines benefits significantly from data availability for model development. However, an effective predictive model is still challenging due to issues with data quality, high-dimensional data, and multi-omics data integration. In this study, we introduce MLG2Net, a deep-learning model inspired by graph neural networks designed to predict DRP in lung cancer cell lines based on pharmacogenomics data. Our model comprises two key components: drug SMILES described by local and global graph networks and cell line genomics are illustrated as a map. Our results show that MLG2Net outperforms three reference graph networks. MLG2Net performance reached a Pearson coefficient correlation ( C C p ) of 0.8616 and a root mean square error (RMSE) of 2.94e-6 in predicting drug responses for Lung Adenocarcinoma (LUAD) cell lines. Subsequent testing on the Lung Squamous Cell Carcinoma (LUSC) dataset reveals lower performance ( C C p : 0.7999, RMSE: 4.08e-6), attributed to the dataset's smaller size influencing model capacity. Moreover, we assessed the model's architecture by isolating its components, with results indicating that the global network is particularly effective in this task. In conclusion, MLG2Net exhibited promising applications in DRP for cancer cell lines, with potential advancements by incorporating larger datasets.

PMID:40208442 | DOI:10.1007/s10916-025-02182-3

  •  
❌