❌

Reading view

Puzzle it Out: Local-to-Global World Model for Offline Multi-Agent Reinforcement Learning

arXiv:2601.07463v1 Announce Type: new Abstract: Offline multi-agent reinforcement learning (MARL) aims to solve cooperative decision-making problems in multi-agent systems using pre-collected datasets. Existing offline MARL methods primarily constrain training within the dataset distribution, resulting in overly conservative policies that struggle to generalize beyond the support of the data. While model-based approaches offer a promising solution by expanding the original dataset with synthetic data generated from a learned world model, the high dimensionality, non-stationarity, and complexity of multi-agent systems make it challenging to accurately estimate the transitions and reward functions in offline MARL. Given the difficulty of directly modeling joint dynamics, we propose a local-to-global (LOGO) world model, a novel framework that leverages local predictions-which are easier to estimate-to infer global state dynamics, thus improving prediction accuracy while implicitly capturing agent-wise dependencies. Using the trained world model, we generate synthetic data to augment the original dataset, expanding the effective state-action space. To ensure reliable policy learning, we further introduce an uncertainty-aware sampling mechanism that adaptively weights synthetic data by prediction uncertainty, reducing approximation error propagation to policies. In contrast to conventional ensemble-based methods, our approach requires only an additional encoder for uncertainty estimation, significantly reducing computational overhead while maintaining accuracy. Extensive experiments across 8 scenarios against 8 baselines demonstrate that our method surpasses state-of-the-art baselines on standard offline MARL benchmarks, establishing a new model-based baseline for generalizable offline multi-agent learning.
  •  

Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation

arXiv:2506.05623v2 Announce Type: replace-cross Abstract: Infrastructure-as-Code (IaC) generation holds significant promise for automating cloud infrastructure provisioning. Recent advances in Large Language Models (LLMs) present a promising opportunity to democratize IaC development by generating deployable infrastructure templates from natural language descriptions. However, current evaluation focuses on syntactic correctness while ignoring deployability, the critical measure of the utility of IaC configuration files. Six state-of-the-art LLMs performed poorly on deployability, achieving only 20.8$\sim$30.2% deployment success rate on the first attempt. In this paper, we construct DPIaC-Eval, the first deployability-centric IaC template benchmark consisting of 153 real-world scenarios cross 58 unique services. Also, we propose an LLM-based deployability-centric framework, dubbed IaCGen, that uses iterative feedback mechanism encompassing format verification, syntax checking, and live deployment stages, thereby closely mirroring the real DevOps workflows. Results show that IaCGen can make 54.6$\sim$91.6% generated IaC templates from all evaluated models deployable in the first 10 iterations. Additionally, human-in-the-loop feedback that provide direct guidance for the deployability errors, can further boost the performance to over 90% passItr@25 on all evaluated LLMs. Furthermore, we explore the trustworthiness of the generated IaC templates on user intent alignment and security compliance. The poor performance (25.2% user requirement coverage and 8.4% security compliance rate) indicates a critical need for continued research in this domain.
  •  

Stereo-seq V2: Spatial mapping of total RNA on FFPE sections with high resolution

Stereo-seq V2 facilitates single-cell-resolution spatial RNA mapping in FFPE samples through random primer capture, uncovering ncRNAs, host-pathogen transcriptome profiling, and spatial immune repertoires in situ.
  •  

OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model

arXiv:2507.05177v3 Announce Type: replace-cross Abstract: Empathetic interaction is a cornerstone of human-machine communication, due to the need for understanding speech enriched with paralinguistic cues and generating emotional and expressive responses. However, the most powerful empathetic LSLMs are increasingly closed off, leaving the crucial details about the architecture, data and development opaque to researchers. Given the critical need for transparent research into the LSLMs and empathetic behavior, we present OpenS2S, a fully open-source, transparent and end-to-end LSLM designed to enable empathetic speech interactions. Based on our empathetic speech-to-text model BLSP-Emo, OpenS2S further employs a streaming interleaved decoding architecture to achieve low-latency speech generation. To facilitate end-to-end training, OpenS2S incorporates an automated data construction pipeline that synthesizes diverse, high-quality empathetic speech dialogues at low cost. By leveraging large language models to generate empathetic content and controllable text-to-speech systems to introduce speaker and emotional variation, we construct a scalable training corpus with rich paralinguistic diversity and minimal human supervision. We release the fully open-source OpenS2S model, including the dataset, model weights, pre-training and fine-tuning codes, to empower the broader research community and accelerate innovation in empathetic speech systems. The project webpage can be accessed at https://casia-lm.github.io/OpenS2S
  •  
  •  

GASPS: A Multi-Omics Framework for Defining Genomic Aberration-Driven Signatures and Predicting Patient Outcomes in Lung Cancer

bioRxiv [Preprint]. 2025 Aug 25:2025.08.21.671519. doi: 10.1101/2025.08.21.671519.

ABSTRACT

Lung cancer is the most common cause of cancer-related death worldwide. Recent advancements in targeted therapies and immunotherapies have achieved remarkable success. However, patient responses to treatments with lung cancer vary substantially. The mutation status of driver genes can direct personalized treatment, but their prognostic value and treatment efficacy are limited. In this study, we developed a statistical framework named Genomic Aberration-Derived Signature for Patient Stratification (GASPS) to characterize the transcriptomic deregulation of driver genomic aberrations and stratify patients. By applying GASPS to The Cancer Genome Atlas Lung Adenocarcinoma (TCGA-LUAD) data, we developed gene signatures for 38 driver genomic aberrations, including gene mutations, amplifications, and deletions. These signatures were applied to independent lung cancer transcriptomic datasets containing a total of 2,226 patient samples. Our results indicated that these driver gene signatures are much more prognostic than their corresponding genomic mutations. Interestingly, the two EGFR-related signatures characterizing EGFR mutation and amplification, respectively, exhibited contrasting associations with prognosis, treatment response, and immune infiltration in the tumor microenvironment. Moreover, the STK11 mutation signature, rather than the mutation status, was found to be predictive of the response and long-term benefit of patients treated with immune checkpoint blockade therapy in lung cancer. This framework is readily applicable to most cancer types using existing data to improve prognostic risk assessment and treatment efficacy by guiding personalized therapies.

PMID:40909579 | PMC:PMC12407784 | DOI:10.1101/2025.08.21.671519

  •  

GASPS: A Multi-Omics Framework for Defining Genomic Aberration-Driven Signatures and Predicting Patient Outcomes in Lung Cancer

bioRxiv [Preprint]. 2025 Aug 25:2025.08.21.671519. doi: 10.1101/2025.08.21.671519.

ABSTRACT

Lung cancer is the most common cause of cancer-related death worldwide. Recent advancements in targeted therapies and immunotherapies have achieved remarkable success. However, patient responses to treatments with lung cancer vary substantially. The mutation status of driver genes can direct personalized treatment, but their prognostic value and treatment efficacy are limited. In this study, we developed a statistical framework named Genomic Aberration-Derived Signature for Patient Stratification (GASPS) to characterize the transcriptomic deregulation of driver genomic aberrations and stratify patients. By applying GASPS to The Cancer Genome Atlas Lung Adenocarcinoma (TCGA-LUAD) data, we developed gene signatures for 38 driver genomic aberrations, including gene mutations, amplifications, and deletions. These signatures were applied to independent lung cancer transcriptomic datasets containing a total of 2,226 patient samples. Our results indicated that these driver gene signatures are much more prognostic than their corresponding genomic mutations. Interestingly, the two EGFR-related signatures characterizing EGFR mutation and amplification, respectively, exhibited contrasting associations with prognosis, treatment response, and immune infiltration in the tumor microenvironment. Moreover, the STK11 mutation signature, rather than the mutation status, was found to be predictive of the response and long-term benefit of patients treated with immune checkpoint blockade therapy in lung cancer. This framework is readily applicable to most cancer types using existing data to improve prognostic risk assessment and treatment efficacy by guiding personalized therapies.

PMID:40909579 | PMC:PMC12407784 | DOI:10.1101/2025.08.21.671519

  •  

Stereo-seq V2: Spatial mapping of total RNA on FFPE sections with high resolution

Cell. 2025 Aug 22:S0092-8674(25)00922-5. doi: 10.1016/j.cell.2025.08.008. Online ahead of print.

ABSTRACT

Performing total RNA profiling on formalin-fixed, paraffin-embedded (FFPE) samples, the predominant sample conservation method in clinical practice, remains challenging for current spatial transcriptomics techniques. Here, we introduce Stereo-seq V2, which employs random primers to capture and sequence RNAs in situ on FFPE sections and provides single-cell resolution. The random-priming-based strategy offers unbiased transcript capturing and uniform gene body coverage, which increase the sensitivity to marker genes, the efficiency of non-polyadenylation (poly(A)) RNA profiling, and immune repertoire coverage. We demonstrated the robust performance of Stereo-seq V2 on clinical FFPE samples using triple-negative breast cancer (TNBC) sections and identified tumor-specific alternative splicing events. In a Mycobacterium tuberculosis (Mtb)-infected mouse model, we monitored gene expression dynamics of host and pathogen transcriptomes simultaneously by utilizing Stereo-seq V2. We also assembled immune repertoires and identified Mtb-specific BCR clones, which could also be observed in human tuberculous lung samples. These results highlight Stereo-seq V2's potential in biomedical research and personalized medicine.

PMID:40882628 | DOI:10.1016/j.cell.2025.08.008

  •  

Development and validation of an integrative 54 biomarker-based risk identification model for multi-cancer in 42,666 individuals: a population-based prospective study to guide advanced screening strategies

Biomark Res. 2025 Aug 11;13(1):101. doi: 10.1186/s40364-025-00812-z.

ABSTRACT

BACKGROUND: Early identification of high-risk individuals is crucial for optimizing cancer screening, particularly when considering expensive and invasive methods such as multi-omics technologies and endoscopic procedures. However, developing a robust, practical multi-cancer risk prediction model that integrates diverse, multi-scale data and with proper validation remains a significant challenge.

METHODS: We initialized the FuSion study by recruiting 42,666 participants from Taizhou, China, with a discovery cohort (n = 16,340) and an independent validation cohort (n = 26,308) after exclusion criteria. We integrated multi-scale data from 54 blood-derived biomarkers and 26 epidemiological exposures to develop a risk prediction model for five common cancers, including lung, esophageal, liver, gastric, and colorectal cancer. Employing five supervised machine learning approaches, we used a LASSO-based feature selection strategy to identify the most informative predictors. The model was trained and internally validated in the discovery cohort, externally applied in the validation cohort, and further evaluated through a prospective clinical follow-up to assess cancer events via clinical examinations.

RESULTS: The final model comprising four key biomarkers along with age, sex, and smoking intensity, achieving an AUROC of 0.767 (95% CI: 0.723-0.814) for five-year risk prediction. High-risk individuals (17.19% of the cohort) accounted for 50.42% of incident cancer cases, with a 15.19-fold increased risk compared to the low-risk group. During follow-up of 2,863 high-risk subjects, 9.64% were newly diagnosed with cancer or precancerous lesions. Notably, cancer detection in the high-risk group was 5.02 times higher than in the low-risk group and 1.74 times higher than in the intermediate-risk group. In particular, the incidence of esophageal cancers in the high-risk group was 16.84 times that of the low-risk group.

CONCLUSIONS: This is the first population-based prospective study in a large Chinese cohort that leverage multi-scale data including biomarkers for multi-cancer risk prediction. Our effective risk stratification model not only enhances early cancer detection but also lays the foundation for the targeted application of advanced screening methods, including but not limited to multi-omics technologies and endoscopy. These findings support precision prevention strategies and the optimal allocation of healthcare resources.

PMID:40790537 | PMC:PMC12341305 | DOI:10.1186/s40364-025-00812-z

  •  

A multiomics dataset of paired CT image and plasma cell-free DNA end motif for patients with pulmonary nodules

Sci Data. 2025 Apr 1;12(1):545. doi: 10.1038/s41597-025-04912-1.

ABSTRACT

Diagnosing lung cancer at a curable stage offers the opportunity for a favorable prognosis. The emerging epigenomics analysis on plasma cell-free DNA (cfDNA), including 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) modifications, has acted as a promising approach facilitating the identification of lung cancer. And, integrating 5mC biomarker with chest computed tomography (CT) image features could optimize the diagnosis of lung cancer, exceeding the performance of models built on single feature. However, the clinical applicability of integrated markers might be limited by the potential risk of overfitting due to small sample size. Hence, we prospectively collected peripheral blood sample and the paired chest CT images of 2032 patients with indeterminate pulmonary nodules across 5 centers, and constructed a large-scale, multi-institutional, multiomics database that encompass CT imaging data and plasma cfDNA fragmentomic in 5mC-, 5hmC-enriched regions. To our best knowledge, this dataset is the first radio-epigenomic dataset with the largest sample size, and provides multi-dimensional insights for early diagnosis of lung cancer, facilitating the individuated management for lung cancer.

PMID:40169596 | PMC:PMC11961589 | DOI:10.1038/s41597-025-04912-1

  •  

Multiple time points for detecting circulating tumor DNA to monitor the response to neoadjuvant therapy in breast cancer: a meta-analysis

BMC Cancer. 2025 Jan 22;25(1):115. doi: 10.1186/s12885-025-13526-0.

ABSTRACT

BACKGROUND: Not all breast cancer (BC) patients can benefit from neoadjuvant therapy (NAT). A poor response may result in patients missing the best opportunity for treatment, ultimately leading to a poor prognosis. Thus, to identify an effective predictor that can assess and predict patient response at early time points, we focused on circulating tumor DNA (ctDNA), which is a vital noninvasive liquid biopsy biomarker. We performed a meta-analysis to explore the predictive value of response by monitoring ctDNA at four time points of NAT using pathologic complete response (pCR) and residual cancer burden (RCB).

METHODS: By searching Embase, PubMed, the Cochrane Library, and the Web of Science until December 24, 2023, we selected studies concerning the relationship between ctDNA and response or prognosis. We analysed the results at the following various time points: baseline (T0), first cycle of NAT (T1), mid-treatment (MT), and end of NAT (EOT). pCR and RCB were used to evaluate the response as the primary endpoint. The secondary endpoint was to investigate the relationship between ctDNA and prognosis. Odds ratios (ORs) and hazard ratios (HRs) were used as effect indicators.

RESULTS: Thirteen reports from twelve studies were eligible for inclusion in this meta-analysis. The results demonstrated that ctDNA negativity was associated with pCR at T1 (OR = 0.34; 95% CI: 0.21-0.57), MT (OR = 0.35; 95% CI: 0.20-0.60), and EOT (OR = 0.38; 95% CI: 0.22-0.66). When RCB was used to evaluate responses, ctDNA negativity was associated with RCB-0/I at the MT (OR = 0.34; 95% CI: 0.21-0.55) and EOT (OR = 0.26; 95% CI: 0.15-0.46). Furthermore, ctDNA positivity at T1 predicted a worse prognosis for patients (HR = 2.73; 95% CI: 1.29-5.75). We also performed a subgroup analysis to more accurately assess the predictive value of ctDNA for triple-negative breast cancer.

CONCLUSIONS: Our meta-analysis suggested that the ctDNA status at the early stage of NAT can predict patient response, which provides evidence for adjusting personalized treatment strategies and improving patient survival.

PROSPERO REGISTRATION NUMBER: CRD42024496465.

PMID:39844103 | PMC:PMC11752932 | DOI:10.1186/s12885-025-13526-0

  •  

Identifying specific functional roles for senescence across cell types

A dual recombinase-mediated genetic system for cell-type-specific lineage tracing, ablation, and gene manipulation of senescent cells reveals distinct roles of senescence across cell types.
  •  

Integrated 5-hydroxymethylcytosine and fragmentation signatures as enhanced biomarkers in lung cancer

Lung cancer is one of most common cancers worldwide, with a 5-year survival rate of less than 20%, which is mainly due to late-stage diagnosis. Noninvasive methods using 5-hydroxymethylation of cytosine (5hmC)...
  •  
❌