❌

Normal view

Explainable multi-omics modeling for risk stratification in pancreatic ductal adenocarcinoma

Gland Surg. 2026 Apr 30;15(4):91. doi: 10.21037/gs-2025-396. Epub 2026 Mar 27.

ABSTRACT

BACKGROUND: Pancreatic ductal adenocarcinoma (PDAC) remains one of the most lethal malignancies due to a lack of reliable tools for individualized risk stratification. A comprehensive understanding of the multi-omics landscape may uncover clinically applicable biomarkers and inform precision prognostic assessment. This study aims to establish a prognostic model directly from the complete omics landscape and extract biomarkers.

METHODS: We developed prognostic models using multi-omics data from a PDAC proteogenomic cohort comprising 75 deceased tumor samples. An independent cohort of 63 deceased PDAC cases from The Cancer Genome Atlas (TCGA)-pancreatic adenocarcinoma (PAAD) was used for external validation. Logistic regression models with least absolute shrinkage and selection operator (LASSO) regularization were constructed, and SHapley Additive exPlanations (SHAP) were applied to evaluate feature importance and identify signature genes. Model selection was based on the average area under the receiver operating characteristic curve (AUROC) across cross-validation folds. Functional validation was performed in PANC-1 cells by knockdown (KD) or overexpression (OE) of representative microRNA-, RNA-, and proteomics-derived signature genes, followed by Cell Counting Kit-8 (CCK-8) proliferation and Transwell migration assays.

RESULTS: Systematic evaluation of 120 multi-omics combinations identified a top-performing prognostic model integrating RNA, microRNA, proteomics, and mutation features. This model achieved a mean AUROC of 0.92Β±0.11 and accuracy of 0.87Β±0.01 on internal validation, and 0.99Β±0.00 and 0.98Β±0.01 on the TCGA test set. The sensitivity, specificity, precision, recall and F1 scores on the TCGA test set were 0.98Β±0.01, 0.97Β±0.02, 0.98Β±0.02, 0.98Β±0.01, 0.98Β±0.01, respectively. SHAP analysis revealed interpretable and clinically relevant prognostic biomarkers, many of which are implicated in immune signaling, metabolic regulation, and cell cycle control. Importantly, modulation of representative signature genes in PANC-1 cells significantly altered proliferation and migration in directions consistent with model-predicted risk associations.

CONCLUSIONS: Our findings demonstrate that explainable multi-omics machine learning frameworks can identify robust prognostic biomarkers and achieve highly accurate survival prediction in PDAC. Functional validation further supports the biological relevance of these signatures, underscoring their translational potential for personalized risk assessment.

PMID:42164702 | PMC:PMC13184197 | DOI:10.21037/gs-2025-396

PRISM: Prompt-Refined In-Context System Modelling for Financial Retrieval

arXiv:2511.14130v2 Announce Type: replace Abstract: With the rapid progress of large language models (LLMs), financial information retrieval has become a critical industrial application. Extracting task-relevant information from lengthy financial filings is essential for both operational and analytical decision-making. We present PRISM, a training-free framework that integrates refined system prompting, in-context learning (ICL), and lightweight multi-agent coordination for document and chunk ranking tasks. Our primary contribution is a systematic empirical study of when each component provides value: prompt engineering delivers consistent performance with minimal overhead, ICL enhances reasoning for complex queries when applied selectively, and multi-agent systems show potential primarily with larger models and careful architectural design. Extensive ablation studies across FinAgentBench, FiQA-2018, and FinanceBench reveal that simpler configurations often outperform complex multi-agent pipelines, providing practical guidance for practitioners. Our best configuration achieves an NDCG@5 of 0.71818 on FinAgentBench, ranking third while being the only training-free approach in the top three. We provide comprehensive feasibility analyses covering latency, token usage, and cost trade-offs to support deployment decisions. The source code is released at https://bit.ly/prism-ailens.

AI-BAAM: AI-Driven Bank Statement Analytics as Alternative Data for Malaysian MSME Credit Scoring

arXiv:2510.16066v4 Announce Type: replace-cross Abstract: Despite accounting for 96.1% of all businesses in Malaysia, access to financing remains one of the most persistent challenges faced by Micro, Small, and Medium Enterprises (MSMEs). Newly established businesses are often excluded from formal credit markets as traditional underwriting approaches rely heavily on credit bureau data. This study investigates the potential of bank statement data as an alternative data source for credit assessment to promote financial inclusion in emerging markets. First, we propose a cash flow-based underwriting pipeline where we utilize bank statement data for end-to-end data extraction and machine learning credit scoring. Second, we introduce a novel dataset of 611 loan applicants from a Malaysian consulting firm. Third, we develop and evaluate credit scoring models based on application information and bank transaction-derived features. Empirical results demonstrate that incorporating bank statement features yields substantial improvements, with our best model achieving an AUROC of 0.806 on validation set, representing a 24.6% improvement over models using application information only. Finally, we will release the anonymized bank transaction dataset to facilitate further research on MSME financial inclusion within Malaysia's emerging economy.
❌