CancerLLM: a large language model in cancer domain
npj Digital Medicine, Published online: 20 February 2026; doi:10.1038/s41746-026-02441-8
CancerLLM: a large language model in cancer domainnpj Digital Medicine, Published online: 20 February 2026; doi:10.1038/s41746-026-02441-8
CancerLLM: a large language model in cancer domainnpj Digital Medicine, Published online: 27 December 2025; doi:10.1038/s41746-025-02253-2
Context matching is not reasoning when performing generalized clinical evaluation of generative language modelsWorld J Hepatol. 2025 Feb 27;17(2):101201. doi: 10.4254/wjh.v17.i2.101201.
ABSTRACT
Liver cancer, particularly hepatocellular carcinoma (HCC), remains a significant global health challenge due to its high mortality rate and late-stage diagnosis. The discovery of reliable biomarkers is crucial for improving early detection and patient outcomes. This review provides a comprehensive overview of current and emerging biomarkers for HCC, including alpha-fetoprotein, des-gamma-carboxy prothrombin, glypican-3, Golgi protein 73, osteopontin, and microRNAs. Despite advancements, the diagnostic limitations of existing biomarkers underscore the urgent need for novel markers that can detect HCC in its early stages. The review emphasizes the importance of integrating multi-omics approaches, combining genomics, proteomics, and metabolomics, to develop more robust biomarker panels. Such integrative methods have the potential to capture the complex molecular landscape of HCC, offering insights into disease mechanisms and identifying targets for personalized therapies. The significance of large-scale validation studies, collaboration between research institutions and clinical settings, and consideration of regulatory pathways for clinical implementation is also discussed. In conclusion, while substantial progress has been made in biomarker discovery, continued research and innovation are essential to address the remaining challenges. The successful translation of these discoveries into clinical practice will require rigorous validation, standardization of protocols, and cross-disciplinary collaboration. By advancing the development and application of novel biomarkers, we can improve the early detection and management of HCC, ultimately enhancing patient survival and quality of life.
PMID:40027561 | PMC:PMC11866143 | DOI:10.4254/wjh.v17.i2.101201
Discov Oncol. 2025 Jan 28;16(1):96. doi: 10.1007/s12672-025-01841-8.
ABSTRACT
BACKGROUND: Pancreatic cancer (PAC) has a complex tumor immune microenvironment, and currently, there is a lack of accurate personalized treatment. Establishing a novel consensus machine learning driven signature (CMLS) that offers a unique predictive model and possible treatment targets for this condition was the goal of this study.
METHODS: This study integrated multiple omics data of PAC patients, applied ten clustering techniques and ten machine learning approaches to construct molecular subtypes for PAC, and created a new CMLS.
RESULTS: Using multi-omics clustering, we discovered two cancer subtypes (CSs) associated with prognosis, among which CS1 exhibited poor prognostic outcomes. Subsequently, 13 central genes were identified through screening, constituting CMLS with a significant prognostic ability. The low CMLS group had a better prognosis and was more likely to possess a "hot" tumor phenotype. The prognosis for the high CMLS group was dismal. Still, the tumor mutation burden (TMB) and tumor neoantigen burden (TNB) levels in this group of patients were higher than in the low CMLS group, which were more favorable for immune therapy response.
CONCLUSION: This study emphasizes that CMLS provides a beneficial instrument for early prediction of patient prognosis and screening of probable patients appropriate for immunotherapy and has broad implications for clinical practice.
PMID:39873820 | PMC:PMC11775367 | DOI:10.1007/s12672-025-01841-8
Int J Med Sci. 2025 Jan 1;22(2):341-356. doi: 10.7150/ijms.104437. eCollection 2025.
ABSTRACT
Background: Gastric cancer (GC) remains a significant global health challenge. This study aimed to comprehensively analyze GC epidemiology and risk factors to inform prevention and intervention strategies. Methods: We analyzed the Global Burden of Disease Study 2021 data, conducted 16 different machine learning (ML) models of NHANES data, performed Mendelian randomization (MR) studies on disease phenotypes, dietary preferences, microbiome, blood-based markers, and integrated differential gene expression and expression quantitative trait loci (eQTL) data from multiple cohorts to identify factors associated with GC risk. Results: Global age-standardized disability-adjusted life year rates (ASDR) for GC declined from 886.24 to 358.42 per 100,000 population between 1990 and 2030, with significant regional disparities. Despite this decline, total disability-adjusted life years show a concerning upward trend from 2015, rising from approximately 22.9 million to a projected 24.3 million by 2030. The slope index of inequality shifted from 87 in 1990 to -184 in 2021, indicating a reversal in GC burden distribution, with higher ASDR now associated with lower socio-demographic index countries. The ML models analysis identified higher levels of clinical characteristics such as phosphorus, calcium, eosinophils percent, and triglycerides, as well as lower levels of iron and monocyte percent, may be associated with an increased risk of GC. MR analyses revealed causal associations between GC risk and disease phenotypes such as Helicobacter pylori infection, chronic gastritis, obesity, depression, and dietary preferences such as dairy and processed meats. Gut microbiome analysis showed associations with microbiome such as Phascolarctobacterium and Ruminococcaceae species. Blood-based markers analysis identified protective and risk effects for cortisol, glutamate, nicotinamide, Natural Killer %lymphocyte, CD4-CD8- T cell Absolute Count, Phosphatidylcholine (16:0_18:1), and Interleukin-1-alpha. Integrated genomic analysis identified 10 genes significantly associated with GC risk, with strong evidence for colocalization in genes such as CCR6 and PILRB. Conclusions: This systematic analysis reveals complex global trends in GC burden and identifies novel clinical, disease phenotypes, dietary preferences, microbial, blood-based, and genetic risk factors. These findings provide potential targets for improved risk stratification, prevention, and intervention strategies to reduce the global burden of GC.
PMID:39781526 | PMC:PMC11704698 | DOI:10.7150/ijms.104437
Nature, Published online: 06 March 2024; doi:10.1038/s41586-024-07148-y
A meta-analysis of genome-wide association studies for 233 circulating metabolites from 33 cohorts reveals more than 400 loci and suggests probable causal genes, providing insights into metabolic pathways and disease aetiology.