❌

Reading view

From Diet to Disease: The Role of the Gut Microbiome and Microbial Metabolites in Horses

Vet Sci. 2026 Aug 18;13(8):823. doi: 10.3390/vetsci13080823.

ABSTRACT

The equine gastrointestinal microbiome plays an essential role in digestion, energy metabolism, immune regulation, and maintenance of intestinal homeostasis. Increasing evidence suggests that microbial-derived metabolites provide a functional link between diet, the microbiome, and host physiology. This narrative review summarizes current knowledge on how dietary factors, including structural carbohydrates, non-structural carbohydrates, protein, and lipids, influence microbial function and metabolite production in horses. Relevant publications available up to June 2026 were identified through searches of PubMed, Web of Science, Scopus, and Google Scholar and were narratively synthesized according to dietary factors, microbial metabolites, associated diseases, and nutritional interventions. Emphasis is placed on major microbial metabolites, including short-chain fatty acids, endotoxins, bile acid derivatives, tryptophan metabolites, and nitrogenous fermentation products, and their potential roles in health and disease. Current evidence supports a central role for the nutrition-microbiota-metabolite axis in gastrointestinal disorders such as colic, colitis, equine gastric ulcer syndrome, and carbohydrate-associated laminitis, whereas its involvement in equine metabolic syndrome and the gut-lung axis remains less well defined. The review also discusses microbiome-targeted nutritional interventions and emerging multi-omics approaches that are improving our understanding of microbial function. Collectively, these findings highlight the potential of microbiome-informed nutritional strategies to support equine health, welfare, disease prevention, and athletic performance.

PMID:42655843 | PMC:PMC13517798 | DOI:10.3390/vetsci13080823

  •  

Towards Multi-Turn Dialog Systems for Industrial Asset Operations and Maintenance

arXiv:2605.24953v1 Announce Type: new Abstract: Industrial asset operations and maintenance question answering is inherently multi-turn, iterative, and highly dependent on external tool invocation. However, the conventional plan-execute single-agent architecture exhibits clear limitations in maintaining cross-turn context, and reusing intermediate results. In this paper, we present a multi-turn dialog system designed for industrial scenarios based on a supervisor-specialist multi-agent architecture. To alleviate tool invocation bottlenecks, the system incorporates structured artifact reuse, dynamic replanning, and parallel tool execution. Evaluation results show that our system achieves better response quality compared with the baseline, with planning effectiveness increasing by 54.5% and task completion improving by 37.8%. System profiling further shows that cross-turn artifact reuse effectively reduces redundant tool invocation, decreasing the tool-time share from 47.3% to 26.3% and making turns 2-5 approximately 4.2x faster than the first turn.
  •  

Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation

arXiv:2605.25036v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) extend large language models with visual understanding, but remain vulnerable to hallucination, where outputs are fluent yet inconsistent with images. Recent studies link this issue to language bias-the tendency of LVLMs to over-rely on text while neglecting visual inputs. Yet most analyses remain empirical without uncovering its underlying cause. In this paper, we provide a systematic study of language bias and identify its root in modality misalignment during training. Our analysis shows that both Visual Instruction Tuning (VIT) and Direct Preference Optimization (DPO) often prioritize textual improvements, which may cause LVLMs to overly lean toward language modeling rather than balanced multimodal understanding. To address this, we propose two simple yet effective methods: Language Bias Regularization (LBR) which mitigates language bias through regularization during instruction tuning, and Language Bias Penalty (LBP), which penalizes language bias in the DPO training process. Extensive experiments across diverse models and benchmarks demonstrate the effectiveness of our approach. LBR consistently improves performance on over ten general benchmarks, while LBP significantly reduces hallucination and improves trustworthiness. Together, these methods not only mitigate language bias but also advance the overall alignment of LVLMs, all without introducing any additional data or auxiliary models. Our code is publicly available at https://github.com/lab-klc/LVLM-Language-Bias.
  •  

Dynamic Dual-Granularity Skill Bank for Agentic RL

arXiv:2603.28716v2 Announce Type: replace Abstract: Agentic RL can benefit substantially from reusable experience, yet existing skill-based methods mainly extract trajectory-level guidance and often lack principled mechanisms for maintaining an evolving skill memory. We propose D2Skill, a dynamic dual-granularity skill bank for agentic RL that organizes reusable experience into task skills for high-level guidance and step skills for fine-grained decision support and error correction. D2Skill jointly trains the policy and skill bank through paired baseline and skill-injected rollouts under the same policy, using their performance gap to derive hindsight utility signals for both skill updating and policy optimization. Built entirely from training-time experience, the skill bank is continuously expanded through reflection and maintained with utility-aware retrieval and pruning. Experiments on ALFWorld, WebShop, and Search-Augmented QA tasks show that D2Skill substantially improves performance over skill-free baselines across models of different scales. Further ablations and analyses show that both dual-granularity skill modeling and dynamic skill maintenance are critical to these gains, while the learned skills exhibit higher utility, transfer across evaluation settings, and introduce only modest training overhead.
  •  

SURGE: Surrogate Gradient Adaptation in Binary Neural Networks

arXiv:2605.10989v3 Announce Type: replace-cross Abstract: The training of Binary Neural Networks (BNNs) is fundamentally based on gradient approximation for non-differentiable binarization operations (e.g., sign function). However, prevailing methods including the Straight-Through Estimator (STE) and its improved variants, rely on hand-crafted designs that suffer from gradient mismatch problem and information loss induced by fixed-range gradient clipping. To address this, we propose SURrogate GradiEnt Adaptation (SURGE), a novel learnable gradient compensation framework with theoretical grounding. SURGE mitigates gradient mismatch through auxiliary backpropagation. Specifically, we design a Dual-Path Gradient Compensator (DPGC) that constructs a parallel full-precision auxiliary branch for each binarized layer, decoupling gradient flow via output decomposition during backpropagation. DPGC enables bias-reduced gradient estimation by leveraging the full-precision branch to estimate components beyond STE's first-order approximation. To further enhance training stability, we introduce an Adaptive Gradient Scaler (AGS) based on an optimal scale factor to dynamically balance inter-branch gradient contributions via norm-based scaling. Experiments on image classification, object detection, and language understanding tasks demonstrate that SURGE performs best over state-of-the-art methods.
  •  

Single-cell spatiotemporal dissection of the human maternal–fetal interface

Nature, Published online: 08 April 2026; doi:10.1038/s41586-026-10316-x

A single-cell multiomic atlas of the human maternal–fetal interface across pregnancy reveals cell types, states and spatial niches, developmental tissue architectures and transcriptional programmes, and identifies cell types with roles in pre-eclampsia, spontaneous preterm birth and miscarriage.
  •  

Efficient amyloid-β degradation in Alzheimer’s disease using SPYTACs

SPYTAC is a synthetic peptide-programmed targeted protein degradation platform harnessing LRP1 to drive lysosomal degradation of extracellular amyloid-β in the brain and periphery. In 5×FAD mice, SPYTAC treatment efficiently degrades amyloid-β, preserves neurons, and improves cognition with reduced neuroinflammation and microhemorrhage when compared with antibody therapy.
  •  

Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning

arXiv:2512.00818v2 Announce Type: replace Abstract: MLLMs MLLMs are beginning to appear in clinical workflows, but their ability to perform complex medical reasoning remains unclear. We present Med-CMR, a fine-grained Medical Complex Multimodal Reasoning benchmark. Med-CMR distinguishes from existing counterparts by three core features: 1) Systematic capability decomposition, splitting medical multimodal reasoning into fine-grained visual understanding and multi-step reasoning to enable targeted evaluation; 2) Challenging task design, with visual understanding across three key dimensions (small-object detection, fine-detail discrimination, spatial understanding) and reasoning covering four clinically relevant scenarios (temporal prediction, causal reasoning, long-tail generalization, multi-source integration); 3) Broad, high-quality data coverage, comprising 20,653 Visual Question Answering (VQA) pairs spanning 11 organ systems and 12 imaging modalities, validated via a rigorous two-stage (human expert + model-assisted) review to ensure clinical authenticity. We evaluate 18 state-of-the-art MLLMs with Med-CMR, revealing GPT-5 as the top-performing commercial model: 57.81 accuracy on multiple-choice questions (MCQs) and a 48.70 open-ended score, outperforming Gemini 2.5 Pro (49.87 MCQ accuracy, 45.98 open-ended score) and leading open-source model Qwen3-VL-235B-A22B (49.34 MCQ accuracy, 42.62 open-ended score). However, specialized medical MLLMs do not reliably outperform strong general models, and long-tail generalization emerges as the dominant failure mode. Med-CMR thus provides a stress test for visual-reasoning integration and rare-case robustness in medical MLLMs, and a rigorous yardstick for future clinical systems.
  •  

Comprehensive multi omics profiling and Mendelian randomization assessment of lipid metabolites in lung cancer prognosis

Discov Oncol. 2026 Mar 23. doi: 10.1007/s12672-026-04893-6. Online ahead of print.

ABSTRACT

BACKGROUND: Lung cancer remains the leading cause of cancer-related mortality worldwide. This study aimed to develop prognostic prediction models for lung squamous cell carcinoma (LUSC) through multi-omics integration using Mendelian randomization analysis.This study addresses a critical gap in lung cancer research through two complementary approaches in major lung cancer subtypes: (1) hypothesis-generating multi-omics analysis in LUSC to identify prognostic biomarkers and characterize the metabolic-immune landscape. This integrated framework provides both predictive tools for personalized medicine and mechanistic insights into metabolic causality.

METHODS: Multi-omics analysis was performed using TCGA data, including RNA-seq, DNA methylation, and whole-exome sequencing. Machine learning models incorporating 15 algorithms were developed and externally validated in two independent GEO cohorts. Mendelian randomization analysis assessed causal relationships between 32 lipid metabolites and SCLC risk. RT-qPCR experiments validated key prognostic genes in lung squamous cell carcinoma (LUSC) cell lines.

RESULTS: The optimal machine learning model (StepCox [forward] + Random Survival Forest) demonstrated superior performance with C-index of 0.73 in internal testing and 0.71 and 0.68 in external validation cohorts. High CD8 + T cell and M1 macrophage infiltration was associated with favorable prognosis. Most lipid metabolites showed no significant causal associations with SCLC risk after multiple testing correction, though two phosphatidylcholine metabolites demonstrated potential protective effects. RT-qPCR validation confirmed significant upregulation of all four key genes in LUSC cell lines.

CONCLUSIONS: This study successfully developed robust machine learning-based prognostic models for LUSC with clinical utility for risk stratification and provided evidence that lipid alterations in lung cancer are likely downstream consequences rather than causal drivers of tumorigenesis.

PMID:41870745 | DOI:10.1007/s12672-026-04893-6

  •  

Comprehensive multi omics profiling and Mendelian randomization assessment of lipid metabolites in lung cancer prognosis

Discov Oncol. 2026 Mar 23. doi: 10.1007/s12672-026-04893-6. Online ahead of print.

ABSTRACT

BACKGROUND: Lung cancer remains the leading cause of cancer-related mortality worldwide. This study aimed to develop prognostic prediction models for lung squamous cell carcinoma (LUSC) through multi-omics integration using Mendelian randomization analysis.This study addresses a critical gap in lung cancer research through two complementary approaches in major lung cancer subtypes: (1) hypothesis-generating multi-omics analysis in LUSC to identify prognostic biomarkers and characterize the metabolic-immune landscape. This integrated framework provides both predictive tools for personalized medicine and mechanistic insights into metabolic causality.

METHODS: Multi-omics analysis was performed using TCGA data, including RNA-seq, DNA methylation, and whole-exome sequencing. Machine learning models incorporating 15 algorithms were developed and externally validated in two independent GEO cohorts. Mendelian randomization analysis assessed causal relationships between 32 lipid metabolites and SCLC risk. RT-qPCR experiments validated key prognostic genes in lung squamous cell carcinoma (LUSC) cell lines.

RESULTS: The optimal machine learning model (StepCox [forward] + Random Survival Forest) demonstrated superior performance with C-index of 0.73 in internal testing and 0.71 and 0.68 in external validation cohorts. High CD8 + T cell and M1 macrophage infiltration was associated with favorable prognosis. Most lipid metabolites showed no significant causal associations with SCLC risk after multiple testing correction, though two phosphatidylcholine metabolites demonstrated potential protective effects. RT-qPCR validation confirmed significant upregulation of all four key genes in LUSC cell lines.

CONCLUSIONS: This study successfully developed robust machine learning-based prognostic models for LUSC with clinical utility for risk stratification and provided evidence that lipid alterations in lung cancer are likely downstream consequences rather than causal drivers of tumorigenesis.

PMID:41870745 | DOI:10.1007/s12672-026-04893-6

  •  

Thinking in Streaming Video

arXiv:2603.12938v1 Announce Type: cross Abstract: Real-time understanding of continuous video streams is essential for interactive assistants and multimodal agents operating in dynamic environments. However, most existing video reasoning approaches follow a batch paradigm that defers reasoning until the full video context is observed, resulting in high latency and growing computational cost that are incompatible with streaming scenarios. In this paper, we introduce ThinkStream, a framework for streaming video reasoning based on a Watch--Think--Speak paradigm that enables models to incrementally update their understanding as new video observations arrive. At each step, the model performs a short reasoning update and decides whether sufficient evidence has accumulated to produce a response. To support long-horizon streaming, we propose Reasoning-Compressed Streaming Memory (RCSM), which treats intermediate reasoning traces as compact semantic memory that replaces outdated visual tokens while preserving essential context. We further train the model using a Streaming Reinforcement Learning with Verifiable Rewards scheme that aligns incremental reasoning and response timing with the requirements of streaming interaction. Experiments on multiple streaming video benchmarks show that ThinkStream significantly outperforms existing online video models while maintaining low latency and memory usage. Code, models and data will be released at https://github.com/johncaged/ThinkStream
  •  

Latilactobacillus curvatus IM01 Alleviates Allergic Airway Inflammation Through Microbial and Metabolic Crosstalk Along the Gut-Lung Axis

Nutrients. 2026 Mar 4;18(5):834. doi: 10.3390/nu18050834.

ABSTRACT

Background: Gut microbiota dysbiosis is critically implicated in the pathogenesis of allergic airway inflammation (AAI) via the gut-lung axis. While Latilactobacillus curvatus is a promising probiotic candidate, its specific immunomodulatory mechanisms in respiratory diseases remain poorly understood. Objective: In this study, we investigated the protective effects and underlying mechanisms of L. curvatus IM01 in an ovalbumin (OVA)-induced murine AAI model using an integrated multi-omics approach. Results: Our results demonstrated that oral administration of L. curvatus IM01 significantly attenuated airway inflammation, suppressed Th2-type immune responses, and reduced serum IgE levels. Crucially, our multi-omics integration revealed a coherent gut-lung axis narrative driven by microbial and metabolic crosstalk. Specifically, 16S rRNA sequencing indicated that L. curvatus IM01 was closely linked to structural shifts in the gut microbial community, notably characterized by an enrichment trend for beneficial genera such as Odoribacter and Lactobacillus. This microbial restructuring was closely associated with a modulated cecal metabolic profile, as untargeted metabolomics exhibited a clear trend toward the restoration of key systemically active immunoregulatory metabolites, including indolelactic acid (ILA) and choline, which have been previously linked to the alleviation of AAI symptoms. Further linking this metabolic shift to respiratory immune tolerance, lung transcriptomic analysis showed that the treatment is strongly associated with the promotion of the differentiation of CD4+ T cells into Foxp3+ regulatory T cells (Tregs). Conclusions: Collectively, these findings suggest a novel potential pathway by which L. curvatus IM01 modulates the gut-lung axis through coordinated microbial and metabolic interventions, highlighting its potential as a therapeutic functional food ingredient for AAI.

PMID:41830004 | PMC:PMC12987261 | DOI:10.3390/nu18050834

  •  

Latilactobacillus curvatus IM01 Alleviates Allergic Airway Inflammation Through Microbial and Metabolic Crosstalk Along the Gut-Lung Axis

Nutrients. 2026 Mar 4;18(5):834. doi: 10.3390/nu18050834.

ABSTRACT

Background: Gut microbiota dysbiosis is critically implicated in the pathogenesis of allergic airway inflammation (AAI) via the gut-lung axis. While Latilactobacillus curvatus is a promising probiotic candidate, its specific immunomodulatory mechanisms in respiratory diseases remain poorly understood. Objective: In this study, we investigated the protective effects and underlying mechanisms of L. curvatus IM01 in an ovalbumin (OVA)-induced murine AAI model using an integrated multi-omics approach. Results: Our results demonstrated that oral administration of L. curvatus IM01 significantly attenuated airway inflammation, suppressed Th2-type immune responses, and reduced serum IgE levels. Crucially, our multi-omics integration revealed a coherent gut-lung axis narrative driven by microbial and metabolic crosstalk. Specifically, 16S rRNA sequencing indicated that L. curvatus IM01 was closely linked to structural shifts in the gut microbial community, notably characterized by an enrichment trend for beneficial genera such as Odoribacter and Lactobacillus. This microbial restructuring was closely associated with a modulated cecal metabolic profile, as untargeted metabolomics exhibited a clear trend toward the restoration of key systemically active immunoregulatory metabolites, including indolelactic acid (ILA) and choline, which have been previously linked to the alleviation of AAI symptoms. Further linking this metabolic shift to respiratory immune tolerance, lung transcriptomic analysis showed that the treatment is strongly associated with the promotion of the differentiation of CD4+ T cells into Foxp3+ regulatory T cells (Tregs). Conclusions: Collectively, these findings suggest a novel potential pathway by which L. curvatus IM01 modulates the gut-lung axis through coordinated microbial and metabolic interventions, highlighting its potential as a therapeutic functional food ingredient for AAI.

PMID:41830004 | PMC:PMC12987261 | DOI:10.3390/nu18050834

  •  

Multi-Omics and Single-Cell Mendelian Randomization Reveal a Potential Role of VNN2 in Lung Adenocarcinoma in Resting Natural Killer Cells

World J Oncol. 2026 Mar 5;17(2):247-255. doi: 10.14740/wjon2689. eCollection 2026 Apr.

ABSTRACT

BACKGROUND: We aimed to evaluate the potential association between genetically predicted vanin-2 (VNN2) expression and lung adenocarcinoma (LUAD) risk, and to explore the immune cell subtype that may underlie this relationship.

METHODS: We integrated whole-blood expression quantitative trait loci (eQTL) data from eQTLGen, plasma protein quantitative trait loci (pQTL) data from deCODE, and LUAD genome-wide association study (GWAS) data from European-ancestry cohorts, together with differential expression analysis using GEPIA2, to identify candidate genes for subsequent single-cell eQTL (sc-eQTL) Mendelian randomization (MR) analysis. For the sc-eQTL analysis, VNN2-associated eQTLs from 14 immune cell types profiled in the OneK1K single-cell eQTL resource were tested for associations with LUAD risk.

RESULTS: Bulk-level MR analysis showed that genetically predicted increases in VNN2 expression and protein levels were significantly associated with a reduced risk of LUAD (eQTL-MR: odds ratio (OR) = 0.964, 95% confidence interval (95% CI), 0.934-0.995; P = 0.024; pQTL-MR: OR = 0.946, 95% CI, 0.921-0.970; P = 2.87 × 10-5). Transcriptomic analyses confirmed significant downregulation of VNN2 in LUAD tumors compared with normal lung tissues. sc-eQTL MR identified the strongest association in resting natural killer (rNK) cells (OR = 0.896, 95% CI, 0.829-0.967; P = 0.005).

CONCLUSIONS: Multi-omics and sc-eQTL MR analyses indicated that genetically predicted increases in VNN2 expression were associated with a reduced risk of LUAD, with the most pronounced effect observed in rNK cells. These findings suggest a potential cell type-specific role of VNN2 in LUAD susceptibility and warrant further studies to validate its biological relevance and clinical implications.

PMID:41822323 | PMC:PMC12978397 | DOI:10.14740/wjon2689

  •  

ViLAM: Distilling Vision-Language Reasoning into Attention Maps for Social Robot Navigation

arXiv:2503.09820v2 Announce Type: replace-cross Abstract: We introduce ViLAM, a novel method for distilling vision-language reasoning from large Vision-Language Models (VLMs) into spatial attention maps for socially compliant robot navigation. Unlike traditional methods that rely on expert demonstrations or human-annotated datasets, ViLAM performs knowledge distillation and fine-tuning at the intermediate layer representation (attention) level by aligning attention maps from a pretrained vision-action model with socially guided attention maps derived from a large VLM. These distilled attention maps highlight key navigational regions in a scene and serve as socially informed spatial cost maps for motion planning. To achieve this, we introduce a novel attention-level distillation loss that fuses knowledge from both sources, generating augmented attention maps with enhanced social awareness. These refined attention maps are then used as a traversability costmap within a socially aware local planner for navigation. We validate our approach through real-world experiments on a Husky wheeled robot, and demonstrate 14.2% - 50% improvements in success rate over existing methods.
  •  

FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference

arXiv:2505.13109v5 Announce Type: replace-cross Abstract: Large language models (LLMs) are widely deployed with rapidly expanding context windows to support increasingly demanding applications. However, long contexts pose significant deployment challenges, primarily due to the KV cache whose size grows proportionally with context length. While KV cache compression methods have been proposed to address this issue, KV dropping methods incur considerable accuracy loss, and KV retrieval methods suffer from significant efficiency bottlenecks. We propose FreeKV, a training-free algorithm-system co-optimization framework to enhance KV retrieval efficiency while preserving accuracy. On the algorithm side, FreeKV introduces speculative retrieval to shift the KV selection and recall processes out of the critical path, combined with fine-grained correction to ensure accuracy. On the system side, FreeKV employs hybrid KV layouts across CPU and GPU memory to eliminate fragmented data transfers, and leverages double-buffered streamed recall to further improve efficiency, enabling effective overlap with computation, full latency hiding, and practical speedups from speculative recall. Experiments demonstrate that FreeKV achieves near-lossless accuracy across various scenarios and models, delivering up to a 13$\times$ speedup compared to SOTA KV retrieval methods. Code is available at https://github.com/sjtu-zhao-lab/FreeKV.
  •  

Mozi: Governed Autonomy for Drug Discovery LLM Agents

arXiv:2603.03655v1 Announce Type: new Abstract: Tool-augmented large language model (LLM) agents promise to unify scientific reasoning with computation, yet their deployment in high-stakes domains like drug discovery is bottlenecked by two critical barriers: unconstrained tool-use governance and poor long-horizon reliability. In dependency-heavy pharmaceutical pipelines, autonomous agents often drift into irreproducible trajectories, where early-stage hallucinations multiplicatively compound into downstream failures. To overcome this, we present Mozi, a dual-layer architecture that bridges the flexibility of generative AI with the deterministic rigor of computational biology. Layer A (Control Plane) establishes a governed supervisor--worker hierarchy that enforces role-based tool isolation, limits execution to constrained action spaces, and drives reflection-based replanning. Layer B (Workflow Plane) operationalizes canonical drug discovery stages -- from Target Identification to Lead Optimization -- as stateful, composable skill graphs. This layer integrates strict data contracts and strategic human-in-the-loop (HITL) checkpoints to safeguard scientific validity at high-uncertainty decision boundaries. Operating on the design principle of ``free-form reasoning for safe tasks, structured execution for long-horizon pipelines,'' Mozi provides built-in robustness mechanisms and trace-level audibility to completely mitigate error accumulation. We evaluate Mozi on PharmaBench, a curated benchmark for biomedical agents, demonstrating superior orchestration accuracy over existing baselines. Furthermore, through end-to-end therapeutic case studies, we demonstrate Mozi's ability to navigate massive chemical spaces, enforce stringent toxicity filters, and generate highly competitive in silico candidates, effectively transforming the LLM from a fragile conversationalist into a reliable, governed co-scientist.
  •  

Unbiased Dynamic Pruning for Efficient Group-Based Policy Optimization

arXiv:2603.04135v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) effectively scales LLM reasoning but incurs prohibitive computational costs due to its extensive group-based sampling requirement. While recent selective data utilization methods can mitigate this overhead, they could induce estimation bias by altering the underlying sampling distribution, compromising theoretical rigor and convergence behavior. To address this limitation, we propose Dynamic Pruning Policy Optimization (DPPO), a framework that enables dynamic pruning while preserving unbiased gradient estimation through importance sampling-based correction. By incorporating mathematically derived rescaling factors, DPPO significantly accelerates GRPO training without altering the optimization objective of the full-batch baseline. Furthermore, to mitigate the data sparsity induced by pruning, we introduce Dense Prompt Packing, a window-based greedy strategy that maximizes valid token density and hardware utilization. Extensive experiments demonstrate that DPPO consistently accelerates training across diverse models and benchmarks. For instance, on Qwen3-4B trained on MATH, DPPO achieves 2.37$\times$ training speedup and outperforms GRPO by 3.36% in average accuracy across six mathematical reasoning benchmarks.
  •  
❌