❌

Reading view

Refined immune-based molecular subtypes of gastric cancer: Integrating mismatch repair status and tumor microenvironment for enhanced immunotherapy prediction

Chin J Cancer Res. 2026 Apr 30;38(2):234-251. doi: 10.21147/j.issn.1000-9604.2026.02.09.

ABSTRACT

OBJECTIVE: Gastric cancer (GC) is heterogeneous, and current mismatch repair (MMR)-based classifications incompletely predict response to immune checkpoint inhibitors (ICIs).

METHODS: RNA sequencing (RNA-seq) and immune infiltration profiles from 189 resected GC were used to derive four refined immune-MMR subtypes (R1-R4) by integrating MMR status, survival, and tumor microenvironment (TME) features. Multi-omics profiling and pathway analysis defined subtype biology. External transcriptomic cohorts and an ICI-treated cohort were classified with Nearest Template Prediction (NTP). Immune response-associated genes were identified from responder vs. non-responder comparisons within the ICI-sensitive subtype and validated by multiplex immunohistochemistry (mIHC).

RESULTS: R1 showed the best prognosis and highest immunotherapy response with objective response rate (ORR) 54.5%, while R4 had the worst prognosis. R2 represented an immune-unresponsive deficient mismatch repair (dMMR) subset, and R3 captured an immune-active proficient mismatch repair (pMMR) subgroup with moderate therapy sensitivity. Multi-omics integration revealed subtype-specific pathways (e.g., ECM remodeling in R1, metabolic reprogramming in R2). Reclassification of pMMR tumors based on transcriptional similarity to R1 identified a New R3 subset with enhanced immune features and higher ICI response. Eight immune response-associated genes (e.g., CXCL10, CXCL11, ELN, GAD1, IL32, MT1E, OR2I1P, SLC3A1) were identified and validated by mIHC for predictive relevance.

CONCLUSIONS: This immune-based molecular framework refines risk stratification beyond conventional MMR categories, identifies ICI-sensitive subsets among both dMMR and pMMR tumors, and proposes candidate biomarkers for patient selection.

PMID:42147371 | PMC:PMC13171420 | DOI:10.21147/j.issn.1000-9604.2026.02.09

  •  

SKILLFOUNDRY: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources

arXiv:2604.03964v1 Announce Type: new Abstract: Modern scientific ecosystems are rich in procedural knowledge across repositories, APIs, scripts, notebooks, documentation, databases, and papers, yet much of this knowledge remains fragmented across heterogeneous artifacts that agents cannot readily operationalize. This gap between abundant scientific know-how and usable agent capabilities is a key bottleneck for building effective scientific agents. We present SkillFoundry, a self-evolving framework that converts such resources into validated agent skills, reusable packages that encode task scope, inputs and outputs, execution steps, environment assumptions, provenance, and tests. SkillFoundry organizes a target domain as a domain knowledge tree, mines resources from high-value branches, extracts operational contracts, compiles them into executable skill packages, and then iteratively expands, repairs, merges, or prunes the resulting library through a closed-loop validation process. SkillFoundry produces a substantially novel and internally valid skill library, with 71.1\% of mined skills differing from existing skill libraries such as SkillHub and SkillSMP. We demonstrate that these mined skills improve coding agent performance on five of the six MoSciBench datasets. We further show that SkillFoundry can design new task-specific skills on demand for concrete scientific objectives, and that the resulting skills substantially improve performance on two challenging genomics tasks: cell type annotation and the scDRS workflow. Together, these results show that automatically mined skills improve agent performance on benchmarks and domain-specific tasks, expand coverage beyond hand-crafted skill libraries, and provide a practical foundation for more capable scientific agents.
  •  

Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based Perspective

arXiv:2502.17262v4 Announce Type: replace-cross Abstract: The escalating scale and cost of Large Language Models (LLMs) training necessitate accurate pre-training prediction of downstream task performance for comprehensive understanding of scaling properties. This is challenged by: 1) the emergence phenomenon, where unpredictable capabilities appearing suddenly at critical model scales; and 2) uneven task difficulty and inconsistent performance scaling patterns, leading to high metric variability. Current prediction methods lack accuracy and reliability. We propose a Clustering-On-Difficulty (COD) framework for downstream performance prediction. The COD framework clusters tasks by their difficulty scaling features, thereby constructing a more stable and predictable task subset that exhibits well-behaved scaling characteristics with the increase of compute budget. We adopt a performance scaling law to predict cluster-wise performance with theoretical support. Predictable subset performance acts as an intermediate predictor for the full evaluation set. We further derive a mapping function to accurately extrapolate the performance of the subset to the full set. Applied to an LLM with 70B parameters, COD achieved a 1.55\% average prediction error across eight key LLM benchmarks, thus providing actionable insights for scaling properties and training monitoring during LLM pre-training.
  •  
❌