❌

Normal view

Subset binding enables detection of multimodal patient subgroup patterns and drug target discovery in idiopathic pulmonary fibrosis

Brief Bioinform. 2026 Mar 1;27(2):bbag153. doi: 10.1093/bib/bbag153.

ABSTRACT

Idiopathic pulmonary fibrosis (IPF) is an intractable lung disease that belongs to idiopathic interstitial pneumonia (IIP) with limited therapeutic options. Conventional patient stratification approaches often fail to integrate diverse data modalities, particularly heterogeneous electronic medical records (EMR) containing mixed discrete and continuous values, with omics data, or fail to extract the interpretable many-to-many relationships crucial for precision medicine. We introduce subset binding (SB), a novel unsupervised algorithm that extends fuzzy association rule mining to robustly integrate heterogeneous clinical data (EMR) and omics data. This framework is uniquely designed to identify clinically meaningful patient subgroup patterns and discover associated molecular signatures based on observable symptoms rather than relying on ambiguous conventional diagnostic categories, such as IIPs. Applying SB to a dataset including 602 samples (from 403 IIPs including IPF patients and 39 healthy controls), we successfully identified 20 proteins linked with key IPF clinical features. Network-based pathway analysis nominated tyrosine kinases as critical drug target candidates, leading to the proposal of ponatinib, a multi-kinase inhibitor, as a candidate therapeutic. Functional validation using a TGF-β-induced epithelial-mesenchymal transition (EMT) model confirmed ponatinib's ability to at least partially suppress TGF-β-induced EMT. This inhibitory effect is consistent with the anti-fibrotic mechanism of the existing IPF drug, nintedanib, and reinforces prior evidence supporting ponatinib's anti-fibrotic property. This study demonstrates that SB enables transparent, reproducible, and robust, molecularly defined patient stratification from multimodal patient data. By establishing a data-driven framework that focuses on observation-based rules, this work lays the critical foundation for future prognostic validation and tailored treatment strategies, offering clinically actionable insights and therapeutic discovery in diagnostically ambiguous diseases like IPF, with ponatinib emerging as a compelling repurposing candidate. Significance statement Idiopathic pulmonary fibrosis (IPF) is a progressive lung disease with limited therapeutic options. IPF is classified as idiopathic interstitial pneumonia (IIP), but distinguishing it from other similar diseases in IIP is not straightforward. The ambiguities in distinguishing IPF from other IIPs necessitate the identification of molecules associated with specific clinical features, rather than relying on solely on diagnosis. Existing methods for multi-omics data analysis often fail to effectively integrate heterogeneous data - such as EMR (containing mixed discrete and continuous values) and omics - or to extract many-to-many molecular-phenotypic relationships. We developed subset binding (SB), a novel, interpretable unsupervised machine learning method to specifically address these technical limitations by integrating EMR and omics data. Our approach successfully detected proteins in serum extracellular vesicles associated with IPF-related features, highlighted several tyrosine kinases as potential drug targets, and proposed the multi-kinase inhibitor ponatinib as a compelling candidate for drug repurposing. This data-driven framework establishes a scalable and interpretable foundation for biomarker and drug target discovery for intractable diseases whose mechanisms are not fully understood.

PMID:41978386 | PMC:PMC13076932 | DOI:10.1093/bib/bbag153

A Joint Neural Baseline for Concept, Assertion, and Relation Extraction from Clinical Text

arXiv:2603.07487v1 Announce Type: cross Abstract: Clinical information extraction (e.g., 2010 i2b2/VA challenge) usually presents tasks of concept recognition, assertion classification, and relation extraction. Jointly modeling the multi-stage tasks in the clinical domain is an underexplored topic. The existing independent task setting (reference inputs given in each stage) makes the joint models not directly comparable to the existing pipeline work. To address these issues, we define a joint task setting and propose a novel end-to-end system to jointly optimize three-stage tasks. We empirically investigate the joint evaluation of our proposal and the pipeline baseline with various embedding techniques: word, contextual, and in-domain contextual embeddings. The proposed joint system substantially outperforms the pipeline baseline by +0.3, +1.4, +3.1 for the concept, assertion, and relation F1. This work bridges joint approaches and clinical information extraction. The proposed approach could serve as a strong joint baseline for future research. The code is publicly available.
❌