❌

Reading view

Self-Improving Pretraining: using post-trained models to pretrain better models

arXiv:2601.21343v3 Announce Type: replace-cross Abstract: Large language models are classically trained in stages: pretraining on raw text followed by post-training for instruction following and reasoning. However, this separation creates a fundamental limitation: many desirable behaviors such as safety, factuality, overall generation quality, and reasoning ability are only added at a late stage, even though the patterns learned earlier strongly shape a model's capabilities. To tackle this issue, we introduce a new way to pretrain and mid-train models that incorporates these behaviors earlier. We utilize an existing strong, post-trained model to both rewrite pretraining data and to judge policy model rollouts, thus using reinforcement earlier in training. In our experiments, we show this can give strong gains in quality, safety, factuality and reasoning.
  •  

Multi-omics integration and machine learning reveal gut-immune signatures in idiopathic pulmonary fibrosis: insights from bulk RNA-seq, single-cell profiles, spatial transcriptomics, and experimental validation

Front Immunol. 2026 Mar 19;17:1730289. doi: 10.3389/fimmu.2026.1730289. eCollection 2026.

ABSTRACT

BACKGROUND: Idiopathic pulmonary fibrosis (IPF) is a progressive, fatal lung disease with limited treatment options and a poor prognosis. Recent studies suggest a critical role for the gut-immune-lung axis in IPF, yet the underlying molecular mechanisms remain unclear.

METHODS: The current study performed in silico multi-omics integration of publicly available datasets, including bulk RNA-seq, single-cell and spatial transcriptomics, as well as peripheral blood multi-omics data to uncover key molecular signatures in IPF. Furthermore, machine learning techniques were utilized to identify core genes, whereas functional analyses and Mendelian randomization were conducted to evaluate the causal relationships among gut microbiota, immune cells, and IPF. Additionally, experimental validation using qPCR and ELISA assays was conducted in vitro, in vivo, and in patient plasma to confirm the expression patterns of key genes.

RESULTS: Across integrated public bulk, single-cell, spatial, and blood multi-omics, CXCL13, IL33, TLR4, and IGF1 were identified as core IPF genes consistently linked to immune infiltration and fibrotic remodeling. Deconvolution, scRNA-seq, and spatial mapping localized their dysregulation to fibroblasts and immune compartments (notably B-cell, macrophage, and mast-cell axes), highlighting fibroblast-immune crosstalk in fibrotic foci. A four-gene model robustly distinguished IPF from controls across cohorts. Mendelian randomization supported a gut-immune-lung axis, indicating causal effects of specific gut taxa on IPF risk via immune phenotypes. qPCR/ELISA in TGF-β1-stimulated fibroblasts, bleomycin mouse lungs, and patient plasma corroborated upregulation of IL33, CXCL13, IGF1 and downregulation of TLR4. Drug-signature reversal nominated cucurbitacin I and temsirolimus; molecular docking was performed as a preliminary in silico, computer-simulation-based assessment of potential ligand-protein interactions between these compounds and the four core targets.

CONCLUSION: This study provides new insights into the importance of gut-immune-lung axis in IPF and identifies CXCL13, IL33, TLR4, and IGF1 as diagnostic signatures and therapeutic targets. By integrating public multi-omics resources with experimental validation, our findings offer a foundation for future diagnostic and treatment strategies aimed at modulating the gut microbiota and immune system in IPF.

PMID:41939867 | PMC:PMC13043422 | DOI:10.3389/fimmu.2026.1730289

  •  

Multi-omics integration and machine learning reveal gut-immune signatures in idiopathic pulmonary fibrosis: insights from bulk RNA-seq, single-cell profiles, spatial transcriptomics, and experimental validation

Front Immunol. 2026 Mar 19;17:1730289. doi: 10.3389/fimmu.2026.1730289. eCollection 2026.

ABSTRACT

BACKGROUND: Idiopathic pulmonary fibrosis (IPF) is a progressive, fatal lung disease with limited treatment options and a poor prognosis. Recent studies suggest a critical role for the gut-immune-lung axis in IPF, yet the underlying molecular mechanisms remain unclear.

METHODS: The current study performed in silico multi-omics integration of publicly available datasets, including bulk RNA-seq, single-cell and spatial transcriptomics, as well as peripheral blood multi-omics data to uncover key molecular signatures in IPF. Furthermore, machine learning techniques were utilized to identify core genes, whereas functional analyses and Mendelian randomization were conducted to evaluate the causal relationships among gut microbiota, immune cells, and IPF. Additionally, experimental validation using qPCR and ELISA assays was conducted in vitro, in vivo, and in patient plasma to confirm the expression patterns of key genes.

RESULTS: Across integrated public bulk, single-cell, spatial, and blood multi-omics, CXCL13, IL33, TLR4, and IGF1 were identified as core IPF genes consistently linked to immune infiltration and fibrotic remodeling. Deconvolution, scRNA-seq, and spatial mapping localized their dysregulation to fibroblasts and immune compartments (notably B-cell, macrophage, and mast-cell axes), highlighting fibroblast-immune crosstalk in fibrotic foci. A four-gene model robustly distinguished IPF from controls across cohorts. Mendelian randomization supported a gut-immune-lung axis, indicating causal effects of specific gut taxa on IPF risk via immune phenotypes. qPCR/ELISA in TGF-β1-stimulated fibroblasts, bleomycin mouse lungs, and patient plasma corroborated upregulation of IL33, CXCL13, IGF1 and downregulation of TLR4. Drug-signature reversal nominated cucurbitacin I and temsirolimus; molecular docking was performed as a preliminary in silico, computer-simulation-based assessment of potential ligand-protein interactions between these compounds and the four core targets.

CONCLUSION: This study provides new insights into the importance of gut-immune-lung axis in IPF and identifies CXCL13, IL33, TLR4, and IGF1 as diagnostic signatures and therapeutic targets. By integrating public multi-omics resources with experimental validation, our findings offer a foundation for future diagnostic and treatment strategies aimed at modulating the gut microbiota and immune system in IPF.

PMID:41939867 | PMC:PMC13043422 | DOI:10.3389/fimmu.2026.1730289

  •  

ST-GDance++: A Scalable Spatial-Temporal Diffusion for Long-Duration Group Choreography

arXiv:2603.22316v1 Announce Type: cross Abstract: Group dance generation from music requires synchronizing multiple dancers while maintaining spatial coordination, making it highly relevant to applications such as film production, gaming, and animation. Recent group dance generation models have achieved promising generation quality, but they remain difficult to deploy in interactive scenarios due to bidirectional attention dependencies. As the number of dancers and the sequence length increase, the attention computation required for aligning music conditions with motion sequences grows quadratically, leading to reduced efficiency and increased risk of motion collisions. Effectively modeling dense spatial-temporal interactions is therefore essential, yet existing methods often struggle to capture such complexity, resulting in limited scalability and unstable multi-dancer coordination. To address these challenges, we propose ST-GDance++, a scalable framework that decouples spatial and temporal dependencies to enable efficient and collision-aware group choreography generation. For spatial modeling, we introduce lightweight distance-aware graph convolutions to capture inter-dancer relationships while reducing computational overhead. For temporal modeling, we design a diffusion noise scheduling strategy together with an efficient temporal-aligned attention mask, enabling stream-based generation for long motion sequences and improving scalability in long-duration scenarios. Experiments on the AIOZ-GDance dataset show that ST-GDance++ achieves competitive generation quality with significantly reduced latency compared to existing methods.
  •  

Low-dose intestinal irradiation enhances the efficacy and prognosis of PD-1 blockade in metastatic non-small cell lung cancer

Clin Cancer Res. 2026 Mar 18. doi: 10.1158/1078-0432.CCR-25-4153. Online ahead of print.

ABSTRACT

PURPOSE: Intestinal low-dose irradiation (ILDR) may enhance immunotherapy efficacy by modulating the gut microbiota and metabolism; however, its role in metastatic non-small cell lung cancer (mNSCLC), particularly in the first-line setting, remains unclear.

EXPERIMENTAL DESIGN: This multicenter retrospective and prospective study included mNSCLC patients receiving first- and second-line programmed cell death protein 1 (PD-1) inhibitors along with abdominopelvic radiotherapy between 2018 and 2025. Patients were stratified by the mean intestinal radiation dose into <1 Gy, 1-3 Gy, and >3 Gy groups and treatment outcomes were compared. The blood and fecal samples were subjected to multi-omics profiling.

RESULTS: g>309 patients were included in the retrospective analysis. Optimal efficacy was observed with a small intestinal mean radiation dose (SIMRD) of 1-3 Gy, showing longer progression-free survival (PFS, 10.2 months) and overall survival (OS, 22.8 months) (P < 0.01), which was consistent across subgroups. Compared with 1-3 Gy, SIMRD >3 Gy (Hazard ratio [HR] = 4.87, P < 0.001) and <1 Gy (HR = 1.85, P < 0.001) independently predicted worse OS. Prospective results confirmed the best disease control rate (P = 0.041) and PFS (P = 0.046) with SIMRD of 1-3 Gy. Responders were enriched in Bacillota, Clostridia, and indole derivatives, particularly indole-3-carboxylic acid. Moreover, the 1-3 Gy group exhibited increased circulating macrophage inflammatory protein-3α and reduced circulating α4β7+ regulatory T cells.

CONCLUSIONS: ILDR influences the efficacy of PD-1 blockade in patients with mNSCLC, particularly when SIMRD is maintained within the 1-3 Gy range, likely through modulation of the gut microbiota-metabolite-immune axis.

PMID:41849236 | DOI:10.1158/1078-0432.CCR-25-4153

  •  

SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents

arXiv:2506.08119v2 Announce Type: replace Abstract: LLM-based agents struggle to execute complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial automation. Existing benchmarks fail to capture the procedural complexity and tool orchestration demands of real-world workflows. We introduce SOP-Bench, a benchmark of 2,000+ tasks from human expert-authored SOPs across 12 business domains (healthcare, logistics, finance, content moderation, etc.). Using a human-AI collaborative framework, experts crafted authentic SOPs while AI generated artifacts (tools, APIs, datasets), all human-validated, yielding realistic tasks with executable interfaces and ground-truth outputs. SOP-Bench serves as a research enabler for systematically investigating agent architectures, model capabilities, and deployment considerations across diverse procedural tasks. We demonstrate its utility through illustrative experiments with a subset of frontier models across Function-Calling (FC) and ReAct agents, revealing critical insights. For example, (1) newer models do not guarantee better performance - Claude 4 family outperforms Claude 4.5 family on ReAct tasks (Claude 4 Opus: 72.4% vs. Claude 4.5 Sonnet: 63.3% task success rate), demonstrating that production upgrades require validation; (2) no single model-agent combination dominates: best performances range from 57% to 100% depending on domain. These examples illustrate how SOP-Bench enables isolating and studying specific dimensions of agent performance without costly production experiments. Our goal is not to rank model capabilities or build optimal agents, but to provide a rigorous evaluation framework that enables the researchers and practitioners to systematically investigate agent design choices, model selection, and deployment strategies. We release the benchmark at https://github.com/amazon-science/sop-bench.
  •  
❌