❌

Reading view

Molecular features of early- vs. late-onset gastric cancer: a systematic review and meta-analysis

BMC Cancer. 2026 Jan 14. doi: 10.1186/s12885-026-15567-5. Online ahead of print.

ABSTRACT

BACKGROUND: Early-onset gastric cancer (EOGC), diagnosed before age 50, is characterized by distinct clinicopathological features, though its molecular landscape remains poorly defined.

METHODS: A systematic literature search of PubMed, Embase, and Web of Science identified studies comparing molecular characteristics of EOGC and late-onset gastric cancer (LOGC). Meta-analyses assessed differences in The Cancer Genome Atlas (TCGA) molecular subtypes, gene mutations, therapeutic biomarkers, and serum tumor markers. Odds ratios (ORs) with 95% confidence intervals (CIs) were calculated; heterogeneity was assessed using the I2 statistic.

RESULTS: EOGC was associated with a higher prevalence of the genomically stable (GS) subtype (OR = 1.71, 95% CI: 1.37-2.12) and a lower prevalence of the chromosomal instability (CIN) subtype (OR = 0.62, 95% CI: 0.50-0.77). CDH1 mutations were more frequent in EOGC (OR = 3.44, 95% CI: 2.85-4.16), while HER2 expression (OR = 0.54, 95% CI: 0.43-0.67), dMMR/MSI-H status (OR = 0.25, 95% CI: 0.12-0.53), and p53 expression (OR = 0.56, 95% CI: 0.39-0.82) were significantly lower. Serum markers including CEA and CA19-9 were also less frequently elevated in EOGC.

CONCLUSION: EOGC represents a biologically distinct subset of gastric cancer with unique genomic and immunological features. These findings support age-specific diagnostic approaches and emphasize the value of multiomic strategies to uncover the mechanisms driving early-onset disease.

PMID:41535782 | DOI:10.1186/s12885-026-15567-5

  •  

Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation

arXiv:2506.05623v2 Announce Type: replace-cross Abstract: Infrastructure-as-Code (IaC) generation holds significant promise for automating cloud infrastructure provisioning. Recent advances in Large Language Models (LLMs) present a promising opportunity to democratize IaC development by generating deployable infrastructure templates from natural language descriptions. However, current evaluation focuses on syntactic correctness while ignoring deployability, the critical measure of the utility of IaC configuration files. Six state-of-the-art LLMs performed poorly on deployability, achieving only 20.8$\sim$30.2% deployment success rate on the first attempt. In this paper, we construct DPIaC-Eval, the first deployability-centric IaC template benchmark consisting of 153 real-world scenarios cross 58 unique services. Also, we propose an LLM-based deployability-centric framework, dubbed IaCGen, that uses iterative feedback mechanism encompassing format verification, syntax checking, and live deployment stages, thereby closely mirroring the real DevOps workflows. Results show that IaCGen can make 54.6$\sim$91.6% generated IaC templates from all evaluated models deployable in the first 10 iterations. Additionally, human-in-the-loop feedback that provide direct guidance for the deployability errors, can further boost the performance to over 90% passItr@25 on all evaluated LLMs. Furthermore, we explore the trustworthiness of the generated IaC templates on user intent alignment and security compliance. The poor performance (25.2% user requirement coverage and 8.4% security compliance rate) indicates a critical need for continued research in this domain.
  •  

Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents

arXiv:2510.24702v1 Announce Type: cross Abstract: Public research results on large-scale supervised finetuning of AI agents remain relatively rare, since the collection of agent training data presents unique challenges. In this work, we argue that the bottleneck is not a lack of underlying data sources, but that a large variety of data is fragmented across heterogeneous formats, tools, and interfaces. To this end, we introduce the agent data protocol (ADP), a light-weight representation language that serves as an "interlingua" between agent datasets in diverse formats and unified agent training pipelines downstream. The design of ADP is expressive enough to capture a large variety of tasks, including API/tool use, browsing, coding, software engineering, and general agentic workflows, while remaining simple to parse and train on without engineering at a per-dataset level. In experiments, we unified a broad collection of 13 existing agent training datasets into ADP format, and converted the standardized ADP data into training-ready formats for multiple agent frameworks. We performed SFT on these data, and demonstrated an average performance gain of ~20% over corresponding base models, and delivers state-of-the-art or near-SOTA performance on standard coding, browsing, tool use, and research benchmarks, without domain-specific tuning. All code and data are released publicly, in the hope that ADP could help lower the barrier to standardized, scalable, and reproducible agent training.
  •  

Integrating deep learning and multi-omics features in radiation pneumonitis prediction for lung cancer patients using PET/CT

BMC Med Imaging. 2025 Oct 27;25(1):426. doi: 10.1186/s12880-025-01971-z.

ABSTRACT

BACKGROUND: To investigate the feasibility and accuracy of PET radiomics features, along with their combination with CT radiomics, dosiomics, and deep learning (DL) features, in predicting radiation pneumonitis (RP) in lung cancer patients treated with volumetric modulated arc therapy (VMAT).

METHODS: A total of 206 and 27 lung cancer patients who underwent VMAT with pre-treatment PET/CT imaging were enrolled from Hospital One and Hospital Two for model training and external validation, respectively. Four machine learning (ML) methods were applied to build radiomics models with features extracted from CT (R_CT), PET (R_PET), radiomics features fused PET/CT (R_fFU) and fused PET/CT images (R_ iFU), as well dosiomics features (D). Three DL models were built to extract features from PET (DL_PET), CT (DL_CT), and fused PET/CT images (DL_FU). The best-performing radiomics and DL models were combined with dosiomics to create the final joint model. ROC curves with AUC, accuracy, sensitivity, and specificity evaluated the performance. A nomogram was constructed using top-performing model features, parameters, and relevant clinical factors.

RESULTS: The extreme gradient boosting (XGBoost) and 18-layer residual neural network (Resnet-18) achieved the best performance. The R+D+DL model combined radiomics, dosiomics, and DL features achieved AUCs of 0.93, 0.92 and 0.89 in the training, internal validaiton and external validation cohorts, respectively. A nomogram constructed with gender, Adaptive RT, SUVp90, and XGBoost-score achieved an AUC of 0.94 for RP prediction in VMAT-treated lung cancer patients using PET/CT.

CONCLUSION: Integrating radiomics, DL, dosiomics features and SUVp90 is promising in the RP prediction for lung cancer patients underwent VMAT using PET/CT images.

PMID:41146084 | DOI:10.1186/s12880-025-01971-z

  •  

High-Sensitive Spatial Proteomics for Pancreatic Cancer Progression Analysis

bioRxiv [Preprint]. 2025 May 5:2025.05.01.651678. doi: 10.1101/2025.05.01.651678.

ABSTRACT

Pancreatic cancer remains as one of the most challenging malignancies to diagnose and treat due to the late development of symptoms and limited early diagnostic options. Intraductal papillary mucinous neoplasms (IPMNs) are non-invasive precursors to invasive pancreatic ductal adenocarcinoma (PDAC)and an understanding of the changes in patterns of protein expression that accompany the progression from normal ductal (ND) cell, to IPMN to PDAC may provide avenues for improved earlier detection. In this study, we present an optimized spatial tissue proteomics workflow, termed SP-Max (Spatial Proteomics Optimized for Maximum Sensitivity and Reproducibility in Minimal Sample), designed to maximize protein recovery and quantification from limited laser micro dissected (LMD) samples. Our workflow enabled the identification of more than 6,000 proteins and the quantification of over 5,200 protein groups from FFPE tissue contours of pancreatic tissues. Comparative analyses across ND, IPMN, and PDAC revealed critical molecular differences in protein pathways and potential markers of progression. SP-Max provides a systematic, reproducible approach that significantly enhances our ability to study precancerous lesions and cancer progression in pancreatic tissues at unprecedented resolution.

PMID:40654937 | PMC:PMC12247709 | DOI:10.1101/2025.05.01.651678

  •  
❌