❌

Normal view

MedDCR: Learning to Design Agentic Workflows for Medical Coding

arXiv:2511.13361v1 Announce Type: new Abstract: Medical coding converts free-text clinical notes into standardized diagnostic and procedural codes, which are essential for billing, hospital operations, and medical research. Unlike ordinary text classification, it requires multi-step reasoning: extracting diagnostic concepts, applying guideline constraints, mapping to hierarchical codebooks, and ensuring cross-document consistency. Recent advances leverage agentic LLMs, but most rely on rigid, manually crafted workflows that fail to capture the nuance and variability of real-world documentation, leaving open the question of how to systematically learn effective workflows. We present MedDCR, a closed-loop framework that treats workflow design as a learning problem. A Designer proposes workflows, a Coder executes them, and a Reflector evaluates predictions and provides constructive feedback, while a memory archive preserves prior designs for reuse and iterative refinement. On benchmark datasets, MedDCR outperforms state-of-the-art baselines and produces interpretable, adaptable workflows that better reflect real coding practice, improving both the reliability and trustworthiness of automated systems.

MolChord: Structure-Sequence Alignment for Protein-Guided Drug Design

arXiv:2510.27671v1 Announce Type: new Abstract: Structure-based drug design (SBDD), which maps target proteins to candidate molecular ligands, is a fundamental task in drug discovery. Effectively aligning protein structural representations with molecular representations, and ensuring alignment between generated drugs and their pharmacological properties, remains a critical challenge. To address these challenges, we propose MolChord, which integrates two key techniques: (1) to align protein and molecule structures with their textual descriptions and sequential representations (e.g., FASTA for proteins and SMILES for molecules), we leverage NatureLM, an autoregressive model unifying text, small molecules, and proteins, as the molecule generator, alongside a diffusion-based structure encoder; and (2) to guide molecules toward desired properties, we curate a property-aware dataset by integrating preference data and refine the alignment process using Direct Preference Optimization (DPO). Experimental results on CrossDocked2020 demonstrate that our approach achieves state-of-the-art performance on key evaluation metrics, highlighting its potential as a practical tool for SBDD.

GASPS: A Multi-Omics Framework for Defining Genomic Aberration-Driven Signatures and Predicting Patient Outcomes in Lung Cancer

bioRxiv [Preprint]. 2025 Aug 25:2025.08.21.671519. doi: 10.1101/2025.08.21.671519.

ABSTRACT

Lung cancer is the most common cause of cancer-related death worldwide. Recent advancements in targeted therapies and immunotherapies have achieved remarkable success. However, patient responses to treatments with lung cancer vary substantially. The mutation status of driver genes can direct personalized treatment, but their prognostic value and treatment efficacy are limited. In this study, we developed a statistical framework named Genomic Aberration-Derived Signature for Patient Stratification (GASPS) to characterize the transcriptomic deregulation of driver genomic aberrations and stratify patients. By applying GASPS to The Cancer Genome Atlas Lung Adenocarcinoma (TCGA-LUAD) data, we developed gene signatures for 38 driver genomic aberrations, including gene mutations, amplifications, and deletions. These signatures were applied to independent lung cancer transcriptomic datasets containing a total of 2,226 patient samples. Our results indicated that these driver gene signatures are much more prognostic than their corresponding genomic mutations. Interestingly, the two EGFR-related signatures characterizing EGFR mutation and amplification, respectively, exhibited contrasting associations with prognosis, treatment response, and immune infiltration in the tumor microenvironment. Moreover, the STK11 mutation signature, rather than the mutation status, was found to be predictive of the response and long-term benefit of patients treated with immune checkpoint blockade therapy in lung cancer. This framework is readily applicable to most cancer types using existing data to improve prognostic risk assessment and treatment efficacy by guiding personalized therapies.

PMID:40909579 | PMC:PMC12407784 | DOI:10.1101/2025.08.21.671519

GASPS: A Multi-Omics Framework for Defining Genomic Aberration-Driven Signatures and Predicting Patient Outcomes in Lung Cancer

bioRxiv [Preprint]. 2025 Aug 25:2025.08.21.671519. doi: 10.1101/2025.08.21.671519.

ABSTRACT

Lung cancer is the most common cause of cancer-related death worldwide. Recent advancements in targeted therapies and immunotherapies have achieved remarkable success. However, patient responses to treatments with lung cancer vary substantially. The mutation status of driver genes can direct personalized treatment, but their prognostic value and treatment efficacy are limited. In this study, we developed a statistical framework named Genomic Aberration-Derived Signature for Patient Stratification (GASPS) to characterize the transcriptomic deregulation of driver genomic aberrations and stratify patients. By applying GASPS to The Cancer Genome Atlas Lung Adenocarcinoma (TCGA-LUAD) data, we developed gene signatures for 38 driver genomic aberrations, including gene mutations, amplifications, and deletions. These signatures were applied to independent lung cancer transcriptomic datasets containing a total of 2,226 patient samples. Our results indicated that these driver gene signatures are much more prognostic than their corresponding genomic mutations. Interestingly, the two EGFR-related signatures characterizing EGFR mutation and amplification, respectively, exhibited contrasting associations with prognosis, treatment response, and immune infiltration in the tumor microenvironment. Moreover, the STK11 mutation signature, rather than the mutation status, was found to be predictive of the response and long-term benefit of patients treated with immune checkpoint blockade therapy in lung cancer. This framework is readily applicable to most cancer types using existing data to improve prognostic risk assessment and treatment efficacy by guiding personalized therapies.

PMID:40909579 | PMC:PMC12407784 | DOI:10.1101/2025.08.21.671519

Integrative single-cell multi-omics profiling of human pancreatic islets identifies T1D-associated genes and regulatory signals

Cell Rep. 2025 Jul 29;44(8):116065. doi: 10.1016/j.celrep.2025.116065. Online ahead of print.

ABSTRACT

Genome-wide association studies (GWASs) have identified over 100 signals associated with type 1 diabetes (T1D). However, it has been challenging to translate any given T1D GWAS signal into mechanistic insights, such as causal variants, their target genes, and the specific cell types involved. Here, we present a comprehensive multi-omic integrative analysis of single-cell/nucleus resolution profiles of gene expression and chromatin accessibility in human pancreatic islets under baseline and T1D-stimulating conditions. We nominate effector cell types for all T1D GWAS signals and the regulatory elements and genes for three independent T1D signals acting through β cells at the DLK1/MEG3, RASGRP1, and TOX loci. Subsequently, we validated the functional impact of these genes and regulatory regions using isogenic human embryonic stem cells (hESCs). We found that loss of RASGRP1 or DLK1, as well as disruption of their corresponding regulatory regions, led to increased β cell apoptosis. Furthermore, β cells derived from isogenic hESCs carrying the T1D risk allele of rs3783355 associated with DLK1 showed elevated β cell death. Through additional RNA sequencing (RNA-seq) and assay for transposase-accessible chromatin using sequencing (ATAC-seq) analyses, we identified five genes upregulated in both RASGRP1-/- and DLK1-/- β-like cells, four of which are near T1D GWAS signals. This integrative approach combining single-cell multi-omics, GWASs, and isogenic human pluripotent stem cell (hPSC)-derived β-like cells illuminates cell type context, genes, single nucleotide polymorphisms (SNPs), and regulatory elements underlying T1D-associated signals, providing insights into the biological functions and molecular mechanisms involved.

PMID:40737125 | DOI:10.1016/j.celrep.2025.116065

Fatty acid-binding proteins in cancers

Int J Surg. 2025 Jul 15. doi: 10.1097/JS9.0000000000003049. Online ahead of print.

ABSTRACT

Fatty acid-binding proteins (FABPs) are intracellular lipid chaperones with molecular weights of approximately 14-15 kDa. By binding and transporting fatty acids and lipid-related molecules, FABPs precisely regulate metabolic pathways, signal transduction, and gene expression, playing a central role in cancer initiation and progression. The 11 identified subtypes (FABP1-FABP12; FABP11 is identical to FABP3) exhibit tissue-specific expression and influence tumor progression through metabolic reprogramming, immune microenvironment modulation, and therapy resistance. Metabolically, FABPs enhance fatty acid uptake, β-oxidation, and synthesis, meeting the high proliferative demands of tumors. In immune regulation, FABP4+ macrophages secrete IL-6 to suppress T cell activity, while FABP6 downregulates MHC-I molecule expression to reduce CD8+ T cell infiltration, fostering an immunosuppressive microenvironment. Regarding therapy resistance, FABP4 enhances mitochondrial β-oxidation to reduce apoptosis in ovarian cancer, and FABP5 promotes chemoresistance in HCC via the HIF-1α pathway. Functional heterogeneity exists among subtypes: FABP7 drives glioblastoma stem cell migration via RXRα signaling, while FABP5 exhibits context-dependent roles, promoting HCC progression but suppressing colorectal cancer (CRC) through mTOR-mediated autophagy. Clinically, FABPs serve as diagnostic biomarkers and therapeutic targets. However, challenges such as insufficient target specificity, cross-cancer heterogeneity, and normal tissue toxicity remain. Future studies should integrate multi-omics and single-cell technologies to elucidate cell-specific mechanisms and develop precise combination therapies for clinical translation.

PMID:40717587 | DOI:10.1097/JS9.0000000000003049

A data-intelligence-intensive bioinformatics copilot system for large-scale omics research and scientific insights

Brief Bioinform. 2025 Jul 2;26(4):bbaf312. doi: 10.1093/bib/bbaf312.

ABSTRACT

Advancements in high-throughput sequencing technologies and artificial intelligence (AI) offer unprecedented opportunities for groundbreaking discoveries in bioinformatics research. However, the challenges of exponential growth of omics data and the rapid development of AI technologies require automated big biological data analysis capability and interdisciplinary knowledge-driven scientific insight. Here, we propose a data-intelligence-intensive bioinformatics copilot (Bio-Copilot) system that synergizes AI capabilities with human researchers to facilitate hypothesis-free exploratory research and inspire novel scientific insights in large-scale omics studies. Bio-Copilot forms high-quality intensive intelligence through close collaboration between multiple agents, driven by large language models (LLMs), and human researchers. To augment the capabilities of Bio-Copilot, this study devises an agent group management strategy, an effective human-agent interaction mechanism, a shared interdisciplinary knowledge database, and continuous learning strategies for the agents. We comprehensively compare Bio-Copilot against GPT-4o and several leading AI agents across diverse bioinformatics tasks, using a broad range of evaluation metrics. Bio-Copilot achieves overall state-of-the-art performance across all tasks, while showcasing exceptional task completeness. Furthermore, on application to constructing a large-scale human lung cell atlas, Bio-Copilot not only reproduces the intricate data integration process detailed in a seminal study but also introduces a recursive, multilevel annotation strategy to capture the continuous nature of cellular states and uncovers the characteristics of rare cell types, highlighting its potential to unravel hidden complexities in biological systems. Beyond the technical achievements, this study also underscores the profound implications of integrating AI capabilities with expert knowledge in accelerating impactful biological discoveries and exploring uncharted territories.

PMID:40639418 | PMC:PMC12245162 | DOI:10.1093/bib/bbaf312

Advancements in liquid biopsy for breast Cancer: Molecular biomarkers and clinical applications

Cancer Treat Rev. 2025 Jun 14;139:102979. doi: 10.1016/j.ctrv.2025.102979. Online ahead of print.

ABSTRACT

Breast cancer is characterized by significant molecular heterogeneity; therefore, there are distinct clinical features, treatment modalities, and prognostic outcomes across its various molecular subtypes. In the era of precision medicine, liquid biopsy has emerged as a convenient and minimally invasive technique capable of dynamically representing the comprehensive tumor gene spectrum. This review systematically elaborates the clinical value of liquid biopsy as a breakthrough tool for precision diagnosis and treatment in breast cancer through dynamic detection of key biomarkers, including circulating tumor DNA (ctDNA), circulating tumor cells (CTCs), exosomes, and non-coding RNA (ncRNA). Specific genetic mutations and methylation signatures in ctDNA can be applied to early breast cancer screening, minimal residual disease monitoring, and tracking drug resistance mechanisms. CTCs enumeration (≥1/7.5 mL in early-stage cancer or ≥ 5/7.5 mL in metastatic cancer) and PD-L1 expression levels demonstrate direct correlations with prognostic stratification and the efficacy of immunotherapy. As the specificity and sensitivity of liquid biopsy continue to improve, personalized treatment strategies, informed by biomarker analysis and targeted precision therapies, have unveiled new avenues of hope for patients with breast cancer. However, several challenges persist in the practical application of liquid biopsy. Despite persistent challenges, such as insufficient standardization and difficulties in resolving low-abundance variants, future advancements should focus on multi-omics integration and AI-driven technological breakthroughs to overcome bottlenecks in clinical translation. This review summarizes cutting-edge liquid biopsy technologies for identifying clinically significant molecular biomarkers, focusing on discussing critical challenges in the strategies to advance precision oncology applications for optimized treatment guidance and disease surveillance in breast cancer.

PMID:40540857 | DOI:10.1016/j.ctrv.2025.102979

Digital phenotyping from wearables using AI characterizes psychiatric disorders and identifies genetic associations

Complex disorders require precise strategies for their characterization. AI-based digital phenotypes from biosensors can be used to predict psychiatric disorders and identify GWAS loci.

Spatial epigenome–transcriptome co-profiling of mammalian tissues

Nature, Published online: 15 March 2023; doi:10.1038/s41586-023-05795-1

The authors present two technologies for spatially resolved, genome-wide, joint profiling of the epigenome and transcriptome by cosequencing chromatin accessibility and gene expression, or histone modifications and gene expression on the same tissue section at near-single-cell resolution.
❌