❌

Reading view

DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extraction

arXiv:2509.14507v2 Announce Type: replace Abstract: Natural Language to SQL (NL2SQL) provides a new model-centric paradigm that simplifies database access for non-technical users by converting natural language queries into SQL commands. Recent advancements, particularly those integrating Retrieval-Augmented Generation (RAG) and Chain-of-Thought (CoT) reasoning, have made significant strides in enhancing NL2SQL performance. However, challenges such as inaccurate task decomposition and keyword extraction by LLMs remain major bottlenecks, often leading to errors in SQL generation. While existing datasets aim to mitigate these issues by fine-tuning models, they struggle with over-fragmentation of tasks and lack of domain-specific keyword annotations, limiting their effectiveness. To address these limitations, we present DeKeyNLU, a novel dataset which contains 1,500 meticulously annotated QA pairs aimed at refining task decomposition and enhancing keyword extraction precision for the RAG pipeline. Fine-tuned with DeKeyNLU, we propose DeKeySQL, a RAG-based NL2SQL pipeline that employs three distinct modules for user question understanding, entity retrieval, and generation to improve SQL generation accuracy. We benchmarked multiple model configurations within DeKeySQL RAG pipeline. Experimental results demonstrate that fine-tuning with DeKeyNLU significantly improves SQL generation accuracy on both BIRD (62.31% to 69.10%) and Spider (84.2% to 88.7%) dev datasets.
  •  

PIVOT: an open-source tool for multi-omic spatial data registration

bioRxiv [Preprint]. 2025 Jun 8:2025.06.08.658506. doi: 10.1101/2025.06.08.658506.

ABSTRACT

Advances in spatial profiling have resulted in the generation of multi-omic atlases that span biological scales. In general, multiple workflows are required for image registration, coordinate registration, and spot deconvolution to integrate modalities. To improve the throughput of registration of multi-omic cohorts, we introduce PIVOT, a user-friendly and open-source interface for streamlined nonlinear registration. We demonstrate PIVOT's strengths through registration of three multi-omic datasets, and show comparison of its performance to existing workflows.

PMID:40661390 | PMC:PMC12259011 | DOI:10.1101/2025.06.08.658506

  •  

Spatial Proteomics and Transcriptomics Reveal Early Immune Cell Organization in Pancreatic Intraepithelial Neoplasia

JCI Insight. 2025 Jun 26:e191595. doi: 10.1172/jci.insight.191595. Online ahead of print.

ABSTRACT

Pancreatic ductal adenocarcinoma (PDAC) has a poor survival rate due to late detection. PDAC arises from precursor microscopic lesions, termed pancreatic intraepithelial neoplasia (PanIN), that develop at least a decade before overt disease--this provides an opportunity to intercept PanIN-to-PDAC progression. However, immune interception strategies require full understanding of PanIN and PDAC cellular architecture. Surgical specimens containing PanIN and PDAC lesions from a unique cohort of five treatment-naïve patients with PDAC were surveyed using spatial-omics (proteomic and transcriptomic). Findings were corroborated by spatial proteomics of PanIN and PDAC from tamoxifen-inducible KPC (tiKPC) mice. We uncovered the organization of lymphoid cells into tertiary lymphoid structures (TLSs) adjacent to PanIN lesions. These TLSs lacked CD21+CD23+ B cells compared to more mature TLSs near the PDAC border. PanINs harbored mostly CD4+ T cells with fewer Tregs and exhausted T cells than PDAC. Peri-tumoral space was enriched with naïve CD4+ and central memory T cells. These observations highlight the opportunity to modulate the immune microenvironment in PanINs before immune exclusion and immunosuppression emerge during progression into PDAC.

PMID:40569674 | DOI:10.1172/jci.insight.191595

  •  
❌