❌

Reading view

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data

arXiv:2605.24414v1 Announce Type: new Abstract: We introduce JT-Safe-V2, a large language model designed to advance the safety and trustworthiness of foundation models, extending our previous JT-Safe model toward a more comprehensive safety-by-design paradigm. JT-Safe-V2 emphasizes the joint optimization of general intelligence and safety-by-design through several key innovations: enriching pre-training data with contextual world knowledge, high-certainty pre-training procedures, and safety strengthening post-training mechanisms for enterprise-oriented agentic capabilities. Building on these safety-enhanced foundation models, we propose Safe-MoMA (Safe Mixture of Models and Agents), a framework that enables traceable and efficient inference through the orchestrated deployment of multiple models and agents. Extensive evaluations demonstrate that JT-Safe-V2 achieves state-of-the-art performance across both general intelligence and safety benchmarks. Moreover, Safe-MoMA reduces inference costs by more than 30\% compared to using the largest standalone model baseline while maintaining comparable performance. To facilitate future research on safety-by-design foundation models, we publicly release the post-trained JT-Safe-V2-35B model checkpoint.
  •  

DeGRe: Dense-supervised Generative Reranking for Recommendation

arXiv:2605.25749v1 Announce Type: cross Abstract: In multi-stage recommender systems, reranking optimizes overall utility by capturing intra-list contextual dependencies, yet its central challenge lies in exploring optimal sequences within an exponentially large permutation space. Recent studies have shifted towards end-to-end generative frameworks, which typically leverage list-wise rewards or preference alignment to guide generator training. However, these methods still face two critical issues. First is the heuristic label bias. Existing methods often construct training targets based on simple rules, such as promoting clicked items to the top, while ignoring causal dependencies within the list context. Second is the credit assignment problem. Sparse list-level posterior rewards fail to directly guide intermediate steps in sequence generation, leading to ambiguous optimization directions. To address these issues, we propose DeGRe (Dense-supervised Generative Reranking), a generative reranking framework that bridges the gap between offline exploration and online efficiency through dense supervision. The core of DeGRe lies in its offline-online decoupled design. During the offline phase, we introduce a Lookahead Evaluator based on cumulative regression, which leverages beam search to actively mine high-value lookahead sequences in the unexposed space. During training, we transform the step-wise value estimations from the evaluator into dense supervision signals and distill them into a lightweight Online Generator. This mechanism enables the generator to internalize lookahead planning capabilities, requiring only a single efficient greedy decoding pass during online inference to approximate the global optimum. Experiments demonstrate that DeGRe outperforms baseline models on public benchmarks and industrial datasets. We have successfully deployed DeGRe on Taobao Flash Shopping, significantly improving online recommendations.
  •  

Explainable multi-omics modeling for risk stratification in pancreatic ductal adenocarcinoma

Gland Surg. 2026 Apr 30;15(4):91. doi: 10.21037/gs-2025-396. Epub 2026 Mar 27.

ABSTRACT

BACKGROUND: Pancreatic ductal adenocarcinoma (PDAC) remains one of the most lethal malignancies due to a lack of reliable tools for individualized risk stratification. A comprehensive understanding of the multi-omics landscape may uncover clinically applicable biomarkers and inform precision prognostic assessment. This study aims to establish a prognostic model directly from the complete omics landscape and extract biomarkers.

METHODS: We developed prognostic models using multi-omics data from a PDAC proteogenomic cohort comprising 75 deceased tumor samples. An independent cohort of 63 deceased PDAC cases from The Cancer Genome Atlas (TCGA)-pancreatic adenocarcinoma (PAAD) was used for external validation. Logistic regression models with least absolute shrinkage and selection operator (LASSO) regularization were constructed, and SHapley Additive exPlanations (SHAP) were applied to evaluate feature importance and identify signature genes. Model selection was based on the average area under the receiver operating characteristic curve (AUROC) across cross-validation folds. Functional validation was performed in PANC-1 cells by knockdown (KD) or overexpression (OE) of representative microRNA-, RNA-, and proteomics-derived signature genes, followed by Cell Counting Kit-8 (CCK-8) proliferation and Transwell migration assays.

RESULTS: Systematic evaluation of 120 multi-omics combinations identified a top-performing prognostic model integrating RNA, microRNA, proteomics, and mutation features. This model achieved a mean AUROC of 0.92Β±0.11 and accuracy of 0.87Β±0.01 on internal validation, and 0.99Β±0.00 and 0.98Β±0.01 on the TCGA test set. The sensitivity, specificity, precision, recall and F1 scores on the TCGA test set were 0.98Β±0.01, 0.97Β±0.02, 0.98Β±0.02, 0.98Β±0.01, 0.98Β±0.01, respectively. SHAP analysis revealed interpretable and clinically relevant prognostic biomarkers, many of which are implicated in immune signaling, metabolic regulation, and cell cycle control. Importantly, modulation of representative signature genes in PANC-1 cells significantly altered proliferation and migration in directions consistent with model-predicted risk associations.

CONCLUSIONS: Our findings demonstrate that explainable multi-omics machine learning frameworks can identify robust prognostic biomarkers and achieve highly accurate survival prediction in PDAC. Functional validation further supports the biological relevance of these signatures, underscoring their translational potential for personalized risk assessment.

PMID:42164702 | PMC:PMC13184197 | DOI:10.21037/gs-2025-396

  •  

Spatial multi-omics defines cancer-associated fibroblasts subtype gradients driving metabolic support and immune remodeling in pancreatic ductal adenocarcinoma

Cancer Lett. 2026 May 16;653:218585. doi: 10.1016/j.canlet.2026.218585. Online ahead of print.

ABSTRACT

Pancreatic ductal adenocarcinoma is characterized by a fibrotic and metabolically active tumor microenvironment where cancer-associated fibroblasts (CAFs) mediate metabolic crosstalk, extracellular matrix (ECM) remodeling, and immune regulation. However, the metabolic and spatial heterogeneity of CAFs remains incompletely understood. We integrated spatial transcriptomics and spatial metabolomics data from PDAC tissues and performed SpatialGlue-based multimodal clustering to define CAF subtypes. To characterize metabolic communication, we developed an optimal transport (OT)-based metabolic inference framework to quantitatively model metabolite association between CAFs and tumor cells. Subtype-specific features were independently validated using an independent spatial metabolomics cohort and multiplex immunofluorescence (mIHC) staining. Furthermore, these features were correlated with clinical outcomes via TCGA-PAAD deconvolution. Spatial multi-omics integration identified three robust CAF subtypes with distinct signatures. OT analysis revealed differential metabolic interactions: CAF_C0 mediated amino acid/peptide transfer, CAF_C1 was the primary source of lipids, while CAF_C2 exhibited limited metabolic association but stronger immune and ECM signaling activity. Deconvolution confirmed that CAF composition was strongly associated with prognosis; CAF_C2 enrichment predicted poorer survival and gemcitabine resistance, whereas a higher CAF_C0/CAF_C1 balance correlated with improved outcomes. By combining spatial multi-omics with OT-based modeling, this study delineates metabolically and spatially distinct CAF states with clinical relevance. Our findings suggest CAFs act as both metabolic donors and immune-ECM regulators, providing new insights into stromal reprogramming and potential subtype-specific therapeutic targets in PDAC.

PMID:42144098 | DOI:10.1016/j.canlet.2026.218585

  •  
❌