❌

Normal view

Therapeutic Co-targeting of Oxidative Phosphorylation and Pyrimidine Synthesis Restores Gemcitabine Response in Pancreatic Ductal Adenocarcinoma

Transl Res. 2026 Sep 13:S1931-5244(26)00195-7. doi: 10.1016/j.trsl.2026.09.009. Online ahead of print.

ABSTRACT

Gemcitabine resistance remains a major barrier to effective therapy in pancreatic ductal adenocarcinoma (PDAC), and current combination regimens show potential to overcome this resistance. Here, we identify the mitochondrial ribosomal proteins MRPS22 and MRPL3 as key metabolic gatekeepers that maintain mitochondrial OXPHOS and pyrimidine metabolism, thereby promoting pancreatic cancer cell proliferation and chemoresistance. Across independent cohorts, high MRPS22/MRPL3 expression associates with poorer survival. Depletion of either gene in PDAC curtailed cell proliferation and xenograft growth, which might be due to an impaired mitochondria function, including destabilized respiratory super-complex assembly, diminished ATP production, and increased oxidative stress. Multi-omics profiling revealed a broad reduction of central-carbon intermediates and a pronounced blockade of de novo pyrimidine synthesis at the dihydroorotate dehydrogenase (DHODH) node. MRPS22 depletion hampered nucleotide-pool generation, and exogenous deoxynucleotides partially rescued PDAC cell growth when MRPs were knocked down. Pharmacologic OXPHOS inhibition increased gemcitabine sensitivity, whereas gemcitabine-resistant derivatives exhibited heightened OXPHOS activity and upregulated mitochondrial ribosomal programs. Co-targeting OXPHOS (antimycin A) or DHODH (brequinar) with gemcitabine produced Loewe synergy in vitro and suppressed growth of gemcitabine-resistant xenografts without affecting body weight. Collectively, these findings established MRPS22/MRPL3 as translation-level drivers of PDAC metabolic fitness and nominate OXPHOS/DHODH blockade as a rational combination strategy to overcome gemcitabine resistance.

PMID:42732873 | DOI:10.1016/j.trsl.2026.09.009

Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses

arXiv:2609.05736v2 Announce Type: replace Abstract: LLM tool agents can be improved without retraining by modifying the runtime harness around a fixed model: prompts, tool interfaces, middleware, state handling, and recovery logic. We study this setting as resource-bounded harness selection for fixed-model multi-turn tool agents, with the search surface scoped to prompts and tool-boundary middleware: edits are guarded intercepts at the tool boundary, not arbitrary rewriting of agent execution logic. Our optimizer-agnostic protocol reports mean held-out lift, worst-condition lift, repeatability, logged cost diagnostics, and RelLift95(B), a conservative estimate of the held-out gain of the harness selected under budget B. We instantiate the protocol with prompt-only and prompt-plus-middleware optimizers, including PRISM, which clusters failures and routes repairs to prompt, tool-boundary middleware, or joint edit surfaces within a Pareto search. On BFCL multi-round, tau2-Retail, and tau2-Telecom, PRISM obtains mean held-out lifts of 14.2, 14.9, and 10.1 percentage points and positive empirical RelLift95 on all three benchmarks, and a component ablation attributes the margin chiefly to failure-surface routing and the edit-pattern constraint. Across optimizers, the results show that some search procedures can occasionally find large gains but still choose brittle updates, so the reliability of the chosen harness should be reported alongside average held-out lift.

Pan-cancer screening and integrative multi-omics and deep learning reveal the prognostic significance of an IBD-CRC shared host-microbe signature in bladder urothelial carcinoma

Transl Oncol. 2026 Sep 9;73:103020. doi: 10.1016/j.tranon.2026.103020. Online ahead of print.

ABSTRACT

BACKGROUND: The prognostic relevance of inflammatory bowel disease (IBD)-colorectal cancer (CRC) shared host-microbe signatures in non-intestinal epithelial malignancies remains unclear. This study aimed to evaluate the prognostic and biological significance of an IBD-CRC shared host-microbe interactome signature in bladder urothelial carcinoma (BLCA).

METHODS: Gene set variation analysis (GSVA) was used to assess the activity of the IBD-CRC shared signature across The Cancer Genome Atlas (TCGA) pan-cancer solid tumor cohorts, including lung, liver, colorectal, and urinary system tumors. In BLCA, weighted gene co-expression network analysis (WGCNA) and least absolute shrinkage and selection operator (LASSO)-Cox regression were applied to construct a prognostic risk model, which was validated in independent transcriptomic cohorts. An attention-based multiple instance learning (MIL) model was developed to predict the LASSO-derived high- or low-risk group from H&E whole-slide images (WSIs), using TCGA cases for training and internal validation and an independent institutional cohort of 39 BLCA patients for external validation. Molecular subtype, immune infiltration, immunohistochemistry (IHC), machine learning, single nucleotide variation/copy number variation (SNV/CNV), single-cell/spatial transcriptomics, and WSI-based deep learning analyses were integrated to characterize the biological relevance of the signature.

RESULTS: High GSVA scores were significantly associated with poor prognosis in BLCA. The LASSO-derived high-risk group was enriched in basal/squamous molecular features and exhibited an immune-infiltrated but immunosuppressive tumor microenvironment, characterized by increased immunosuppressive cell infiltration and elevated immune checkpoint expression. Conventional IHC markers supported distinct subtype-related protein phenotypes between risk groups. Single-cell and spatial transcriptomic analyses revealed that malignant cells with high signature activity were enriched in Wnt, Hippo, and cell adhesion pathways. The WSI-based MIL model achieved an area under the curve (AUC) of 0.852 in the internal validation cohort. Machine learning and SNV/CNV analyses further characterized key molecular features associated with the LASSO risk score, including AKR1B1, LY6E, MEST, and others. Pan-cancer characterization of AKR1B1 across multiple malignancies, including lung adenocarcinoma (LUAD), liver hepatocellular carcinoma (LIHC), and kidney renal clear cell carcinoma (KIRC), revealed cancer-type-specific associations with immunosuppressive microenvironmental features and tumor stemness.

CONCLUSION: The IBD-CRC shared host-microbe signature has significant prognostic value in BLCA and is associated with basal/squamous differentiation, immunosuppressive microenvironmental features, genomic alteration patterns, and malignant cell functional heterogeneity. The integrated multi-omics framework and externally validated pathology AI model provide potential tools for BLCA risk stratification and biological interpretation.

PMID:42715652 | DOI:10.1016/j.tranon.2026.103020

Pan-cancer screening and integrative multi-omics and deep learning reveal the prognostic significance of an IBD-CRC shared host-microbe signature in bladder urothelial carcinoma

Transl Oncol. 2026 Sep 9;73:103020. doi: 10.1016/j.tranon.2026.103020. Online ahead of print.

ABSTRACT

BACKGROUND: The prognostic relevance of inflammatory bowel disease (IBD)-colorectal cancer (CRC) shared host-microbe signatures in non-intestinal epithelial malignancies remains unclear. This study aimed to evaluate the prognostic and biological significance of an IBD-CRC shared host-microbe interactome signature in bladder urothelial carcinoma (BLCA).

METHODS: Gene set variation analysis (GSVA) was used to assess the activity of the IBD-CRC shared signature across The Cancer Genome Atlas (TCGA) pan-cancer solid tumor cohorts, including lung, liver, colorectal, and urinary system tumors. In BLCA, weighted gene co-expression network analysis (WGCNA) and least absolute shrinkage and selection operator (LASSO)-Cox regression were applied to construct a prognostic risk model, which was validated in independent transcriptomic cohorts. An attention-based multiple instance learning (MIL) model was developed to predict the LASSO-derived high- or low-risk group from H&E whole-slide images (WSIs), using TCGA cases for training and internal validation and an independent institutional cohort of 39 BLCA patients for external validation. Molecular subtype, immune infiltration, immunohistochemistry (IHC), machine learning, single nucleotide variation/copy number variation (SNV/CNV), single-cell/spatial transcriptomics, and WSI-based deep learning analyses were integrated to characterize the biological relevance of the signature.

RESULTS: High GSVA scores were significantly associated with poor prognosis in BLCA. The LASSO-derived high-risk group was enriched in basal/squamous molecular features and exhibited an immune-infiltrated but immunosuppressive tumor microenvironment, characterized by increased immunosuppressive cell infiltration and elevated immune checkpoint expression. Conventional IHC markers supported distinct subtype-related protein phenotypes between risk groups. Single-cell and spatial transcriptomic analyses revealed that malignant cells with high signature activity were enriched in Wnt, Hippo, and cell adhesion pathways. The WSI-based MIL model achieved an area under the curve (AUC) of 0.852 in the internal validation cohort. Machine learning and SNV/CNV analyses further characterized key molecular features associated with the LASSO risk score, including AKR1B1, LY6E, MEST, and others. Pan-cancer characterization of AKR1B1 across multiple malignancies, including lung adenocarcinoma (LUAD), liver hepatocellular carcinoma (LIHC), and kidney renal clear cell carcinoma (KIRC), revealed cancer-type-specific associations with immunosuppressive microenvironmental features and tumor stemness.

CONCLUSION: The IBD-CRC shared host-microbe signature has significant prognostic value in BLCA and is associated with basal/squamous differentiation, immunosuppressive microenvironmental features, genomic alteration patterns, and malignant cell functional heterogeneity. The integrated multi-omics framework and externally validated pathology AI model provide potential tools for BLCA risk stratification and biological interpretation.

PMID:42715652 | DOI:10.1016/j.tranon.2026.103020

Pan-cancer screening and integrative multi-omics and deep learning reveal the prognostic significance of an IBD-CRC shared host-microbe signature in bladder urothelial carcinoma

Transl Oncol. 2026 Sep 9;73:103020. doi: 10.1016/j.tranon.2026.103020. Online ahead of print.

ABSTRACT

BACKGROUND: The prognostic relevance of inflammatory bowel disease (IBD)-colorectal cancer (CRC) shared host-microbe signatures in non-intestinal epithelial malignancies remains unclear. This study aimed to evaluate the prognostic and biological significance of an IBD-CRC shared host-microbe interactome signature in bladder urothelial carcinoma (BLCA).

METHODS: Gene set variation analysis (GSVA) was used to assess the activity of the IBD-CRC shared signature across The Cancer Genome Atlas (TCGA) pan-cancer solid tumor cohorts, including lung, liver, colorectal, and urinary system tumors. In BLCA, weighted gene co-expression network analysis (WGCNA) and least absolute shrinkage and selection operator (LASSO)-Cox regression were applied to construct a prognostic risk model, which was validated in independent transcriptomic cohorts. An attention-based multiple instance learning (MIL) model was developed to predict the LASSO-derived high- or low-risk group from H&E whole-slide images (WSIs), using TCGA cases for training and internal validation and an independent institutional cohort of 39 BLCA patients for external validation. Molecular subtype, immune infiltration, immunohistochemistry (IHC), machine learning, single nucleotide variation/copy number variation (SNV/CNV), single-cell/spatial transcriptomics, and WSI-based deep learning analyses were integrated to characterize the biological relevance of the signature.

RESULTS: High GSVA scores were significantly associated with poor prognosis in BLCA. The LASSO-derived high-risk group was enriched in basal/squamous molecular features and exhibited an immune-infiltrated but immunosuppressive tumor microenvironment, characterized by increased immunosuppressive cell infiltration and elevated immune checkpoint expression. Conventional IHC markers supported distinct subtype-related protein phenotypes between risk groups. Single-cell and spatial transcriptomic analyses revealed that malignant cells with high signature activity were enriched in Wnt, Hippo, and cell adhesion pathways. The WSI-based MIL model achieved an area under the curve (AUC) of 0.852 in the internal validation cohort. Machine learning and SNV/CNV analyses further characterized key molecular features associated with the LASSO risk score, including AKR1B1, LY6E, MEST, and others. Pan-cancer characterization of AKR1B1 across multiple malignancies, including lung adenocarcinoma (LUAD), liver hepatocellular carcinoma (LIHC), and kidney renal clear cell carcinoma (KIRC), revealed cancer-type-specific associations with immunosuppressive microenvironmental features and tumor stemness.

CONCLUSION: The IBD-CRC shared host-microbe signature has significant prognostic value in BLCA and is associated with basal/squamous differentiation, immunosuppressive microenvironmental features, genomic alteration patterns, and malignant cell functional heterogeneity. The integrated multi-omics framework and externally validated pathology AI model provide potential tools for BLCA risk stratification and biological interpretation.

PMID:42715652 | DOI:10.1016/j.tranon.2026.103020

From Prompt Optimization to Multi-Dimensional Credibility Evaluation: Enhancing Trustworthiness of Chinese LLM-Generated Liver MRI Reports -- with Preliminary Extension to Lung Cancer

arXiv:2510.23008v3 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated promising performance in generating diagnostic conclusions from imaging findings, thereby supporting radiology reporting, trainee education, and quality control. However, systematic guidance on how to optimize prompt design across different clinical contexts remains underexplored. Moreover, a comprehensive and standardized framework for assessing the trustworthiness of LLM-generated radiology reports is yet to be established. This study aims to enhance the trustworthiness of LLM-generated liver MRI reports by introducing a Multi-Dimensional Credibility Assessment (MDCA) framework and providing guidance on institution-specific prompt optimization. The proposed framework is applied to evaluate and compare the performance of several advanced LLMs, including Kimi-K2-Instruct-0905, Qwen3-235B-A22B-Instruct-2507, DeepSeek-V3, and ByteDance-Seed-OSS-36B-Instruct, using the SiliconFlow platform.

Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs

arXiv:2601.21463v3 Announce Type: replace-cross Abstract: Existing speech editing detection (SED) datasets are predominantly constructed using manual splicing or limited editing operations, resulting in restricted diversity and poor coverage of realistic editing scenarios. Meanwhile, current SED methods rely heavily on frame-level supervision to detect observable acoustic anomalies, which fundamentally limits their ability to handle deletion-type edits, where the manipulated content is entirely absent from the signal. To address these challenges, we present a unified framework that bridges speech editing detection and content localization through a generative formulation based on Audio Large Language Models (Audio LLMs). We first introduce AiEdit, https://huggingface.co/datasets/JunXueTech/AiEdit, a large-scale bilingual dataset (approximately 140 hours) that covers addition, deletion, and modification operations using state-of-the-art end-to-end speech editing systems, providing a more realistic benchmark for modern threats. Building upon this, we reformulate SED as a structured text generation task, enabling joint reasoning over edit type identification, and content localization. To enhance the grounding of generative models in acoustic evidence, we propose a prior-enhanced prompting strategy that injects word-level probabilistic cues derived from a frame-level detector. Furthermore, we introduce an acoustic consistency-aware loss that explicitly enforces the separation between normal and anomalous acoustic representations in the latent space. Experimental results demonstrate that the proposed approach consistently outperforms existing methods across both detection and localization tasks.

The landscape of virtual simulation in undergraduate dental education with implications for artificial intelligence incorporation

npj Digital Medicine, Published online: 26 May 2026; doi:10.1038/s41746-026-02806-z

The landscape of virtual simulation in undergraduate dental education with implications for artificial intelligence incorporation

Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models

arXiv:2604.03302v1 Announce Type: cross Abstract: While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focus. Despite their improvements, current MLLMs struggle significantly with high-level physics reasoning. In this work, we investigate the first step of physical reasoning, i.e., intuitive physics understanding, revealing substantial limitations in understanding the dynamics of continuum objects. To isolate and evaluate this specific capability, we introduce two fundamental benchmark tasks: Next Frame Selection (NFS) and Temporal Coherence Verification (TCV). Our experiments demonstrate that even state-of-the-art MLLMs perform poorly on these foundational tasks. To address this limitation, we propose Scene Dynamic Field (SDF), a concise approach that leverages physics simulators within a multi-task fine-tuning framework. SDF substantially improves performance, achieving up to 20.7% gains on fluid tasks while showing strong generalization to unseen physical domains. This work not only highlights a critical gap in current MLLMs but also presents a promising cost-efficient approach for developing more physically grounded MLLMs. Our code and data are available at https://github.com/andylinx/Scene-Dynamic-Field.

ROSClaw: A Hierarchical Semantic-Physical Framework for Heterogeneous Multi-Agent Collaboration

arXiv:2604.04664v1 Announce Type: cross Abstract: The integration of large language models (LLMs) with embodied agents has improved high-level reasoning capabilities; however, a critical gap remains between semantic understanding and physical execution. While vision-language-action (VLA) and vision-language-navigation (VLN) systems enable robots to perform manipulation and navigation tasks from natural language instructions, they still struggle with long-horizon sequential and temporally structured tasks. Existing frameworks typically adopt modular pipelines for data collection, skill training, and policy deployment, resulting in high costs in experimental validation and policy optimization. To address these limitations, we propose ROSClaw, an agent framework for heterogeneous robots that integrates policy learning and task execution within a unified vision-language model (VLM) controller. The framework leverages e-URDF representations of heterogeneous robots as physical constraints to construct a sim-to-real topological mapping, enabling real-time access to the physical states of both simulated and real-world agents. We further incorporate a data collection and state accumulation mechanism that stores robot states, multimodal observations, and execution trajectories during real-world execution, enabling subsequent iterative policy optimization. During deployment, a unified agent maintains semantic continuity between reasoning and execution, and dynamically assigns task-specific control to different agents, thereby improving robustness in multi-policy execution. By establishing an autonomous closed-loop framework, ROSClaw minimizes the reliance on robot-specific development workflows. The framework supports hardware-level validation, automated generation of SDK-level control programs, and tool-based execution, enabling rapid cross-platform transfer and continual improvement of robotic skills. Ours project page: https://www.rosclaw.io/.

Predictive Value of Machine Learning for Poststroke Mortality Risk: Systematic Review and Meta-Analysis

Background: People with stroke face a high mortality risk, and an accurate prediction model is essential to the guidance of clinical decision-making in this population. Recently, with growing attention paid to machine learning (ML) in stroke care, some researchers have investigated the effectiveness of ML in predicting the mortality risk in stroke. However, systematic evidence is still lacking for its effectiveness. Objective: This systematic review aims to evaluate the value of ML in predicting the stroke mortality risk. The findings are expected to offer an evidence-based basis for developing and assessing clinical risk prediction tools. Methods: A search was made in Cochrane Library, PubMed, Embase, and Web of Science up to June 23, 2025, and studies that reported a complete performance of ML in predicting stroke mortality were included. Studies with only risk factors analyzed were excluded. The risk of bias of the included studies was assessed using PROBAST (Prediction model Risk of Bias Assessment Tool). Pooled risk ratios with 95% CIs and prediction intervals (PIs) were derived using the Hartung-Knapp-Sidik-Jonkman method under a random-effects model. Subgroup analyses were also conducted by model type, stroke type, patient source, and treatment background. Moreover, a metaregression was conducted on the C-index for out-of-hospital mortality at different time points to explore the influence of time factors on the model’s predictive performance. Results: Sixty-eight studies were included (23 predicting in-hospital mortality and 45 predicting out-of-hospital mortality), describing the development of 75 prediction models and 43 external validations. The follow-up period was 1 month to 15 years. For predicting in-hospital mortality, the external validation set had a pooled C-index of 0.727 (95% CI 0.677-0.781, 95% PI 0.521-1.000), with sensitivity and specificity of 0.64 (95% CI 0.57-0.70) and 0.74 (95% CI 0.70-0.77), respectively. For predicting out-of-hospital mortality, the pooled C-index was 0.847 (95% CI 0.808-0.887, 95% PI 0.750-0.956) in the external validation set, with sensitivity and specificity of 0.71 (95% CI 0.55-0.82) and 0.76 (95% CI 0.74-0.78), respectively. Comparatively, the overall pooled C-indexes were 0.788 (95% CI 0.766-0.810, 95% PI 0.621-0.999) and 0.812 (95% CI 0.798-0.826, 95% PI 0.693-0.952), respectively. The metaregression revealed a gradual decline in the predictive performance of the overall model and logistic regression model alone, whereas a random forest model maintained sustained performance. Age, National Institutes of Health Stroke Scale score, and stroke-related complications were the most frequently used variables for modeling. Conclusions: This is the first meta-analysis to demonstrate that ML-based prediction of stroke mortality is feasible. The performance of ML supports its role as an auxiliary tool for identifying high-risk populations, thereby optimizing clinical monitoring and resource allocation. However, due to substantial heterogeneity and a relatively high risk of bias in available studies, caution is warranted in real-world application. The effectiveness of ML may vary across settings, and external validation is recommended before broader implementation. Trial Registration: PROSPERO CRD420251086321; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251086321

X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving

arXiv:2603.19979v2 Announce Type: replace-cross Abstract: Scalable and reliable evaluation is increasingly critical in the end-to-end era of autonomous driving, where vision--language--action (VLA) policies directly map raw sensor streams to driving actions. Yet, current evaluation pipelines still rely heavily on real-world road testing, which is costly, biased toward limited scenario coverage, and difficult to reproduce. These challenges motivate a real-world simulator that can generate realistic future observations under proposed actions, while remaining controllable and stable over long horizons. We present X-World, an action-conditioned multi-camera generative world model that simulates future observations directly in video space. Given synchronized multi-view camera history and a future action sequence, X-World generates future multi-camera video streams that follow the commanded actions. To ensure reproducible and editable scene rollouts, X-World further supports optional controls over dynamic traffic agents and static road elements, and retains a text-prompt interface for appearance-level control (e.g., weather and time of day). Beyond world simulation, X-World also enables video style transfer by conditioning on appearance prompts while preserving the underlying action and scene dynamics. At the core of X-World is a multi-view latent video generator designed to explicitly encourage cross-view geometric consistency and temporal coherence under diverse control signals. Experiments show that X-World achieves high-quality multi-view video generation with (i) strong view consistency across cameras, (ii) stable temporal dynamics over long rollouts, and (iii) high controllability with strict action following and faithful adherence to optional scene controls. These properties make X-World a practical foundation for scalable and reproducible evaluation.

Gene regulatory landscape dissected by single-cell four-omics sequencing

Nature, Published online: 01 April 2026; doi:10.1038/s41586-026-10322-z

Combining single-cell parallel profiling of genome conformation, histone modifications, chromatin accessibility and gene expression reveals dynamics and intranuclear spatial clustering of epigenome profiles, enabling sophisticated analysis of the regulatory landscape across cell types and tissues.

WiseMind: a knowledge-guided multi-agent framework for accurate and empathetic psychiatric diagnosis

npj Digital Medicine, Published online: 25 March 2026; doi:10.1038/s41746-026-02559-9

WiseMind: a knowledge-guided multi-agent framework for accurate and empathetic psychiatric diagnosis

A Robust Incomplete Multimodal Low-Rank Adaptation Approach for Emotion Recognition

arXiv:2507.11202v1 Announce Type: cross Abstract: Multimodal Emotion Recognition (MER) often encounters incomplete multimodality in practical applications due to sensor failures or privacy protection requirements. While existing methods attempt to address various incomplete multimodal scenarios by balancing the training of each modality combination through additional gradients, these approaches face a critical limitation: training gradients from different modality combinations conflict with each other, ultimately degrading the performance of the final prediction model. In this paper, we propose a unimodal decoupled dynamic low-rank adaptation method based on modality combinations, named MCULoRA, which is a novel framework for the parameter-efficient training of incomplete multimodal learning models. MCULoRA consists of two key modules, modality combination aware low-rank adaptation (MCLA) and dynamic parameter fine-tuning (DPFT). The MCLA module effectively decouples the shared information from the distinct characteristics of individual modality combinations. The DPFT module adjusts the training ratio of modality combinations based on the separability of each modality's representation space, optimizing the learning efficiency across different modality combinations. Our extensive experimental evaluation in multiple benchmark datasets demonstrates that MCULoRA substantially outperforms previous incomplete multimodal learning approaches in downstream task accuracy.

More Bang for the Buck: Process Reward Modeling with Entropy-Driven Uncertainty

arXiv:2503.22233v4 Announce Type: replace-cross Abstract: We introduce the Entropy-Driven Uncertainty Process Reward Model (EDU-PRM), a novel entropy-driven training framework for process reward modeling that enables dynamic, uncertainty-aligned segmentation of complex reasoning steps, eliminating the need for costly manual step annotations. Unlike previous Process Reward Models (PRMs) that rely on static partitioning and human labeling, EDU-PRM automatically anchors step boundaries at tokens with high predictive entropy, effectively capturing intrinsic logical transitions and facilitating efficient exploration of diverse reasoning paths. On the ProcessBench benchmark, EDU-PRM outperforms strong public PRM baselines, such as Math-Shepherd PRM and Omega PRM, and EDU-PRM achieves comparable results with SOTA models while only using 1.5% training data. Furthermore, by leveraging our proposed EDU sampling strategy, we observe accuracy boosts from 64.7% to 67.3% for generative reasoning tasks, accompanied by a reduction of 32% in token usage. These findings underscore the potential of EDU-PRM as a scalable and annotation-efficient paradigm for process supervision in mathematical reasoning, paving the way for more efficient and robust approaches to complex mathematical problem solving.

Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters

arXiv:2602.10604v2 Announce Type: replace-cross Abstract: We introduce Step 3.5 Flash, a sparse Mixture-of-Experts (MoE) model that bridges frontier-level agentic intelligence and computational efficiency. We focus on what matters most when building agents: sharp reasoning and fast, reliable execution. Step 3.5 Flash pairs a 196B-parameter foundation with 11B active parameters for efficient inference. It is optimized with interleaved 3:1 sliding-window/full attention and Multi-Token Prediction (MTP-3) to reduce the latency and cost of multi-round agentic interactions. To reach frontier-level intelligence, we design a scalable reinforcement learning framework that combines verifiable signals with preference feedback, while remaining stable under large-scale off-policy training, enabling consistent self-improvement across mathematics, code, and tool use. Step 3.5 Flash demonstrates strong performance across agent, coding, and math tasks, achieving 85.4% on IMO-AnswerBench, 86.4% on LiveCodeBench-v6 (2024.08-2025.05), 88.2% on tau2-Bench, 69.0% on BrowseComp (with context management), and 51.0% on Terminal-Bench 2.0, comparable to frontier models such as GPT-5.2 xHigh and Gemini 3.0 Pro. By redefining the efficiency frontier, Step 3.5 Flash provides a high-density foundation for deploying sophisticated agents in real-world industrial environments.

Rethinking Diffusion Models with Symmetries through Canonicalization with Applications to Molecular Graph Generation

arXiv:2602.15022v1 Announce Type: cross Abstract: Many generative tasks in chemistry and science involve distributions invariant to group symmetries (e.g., permutation and rotation). A common strategy enforces invariance and equivariance through architectural constraints such as equivariant denoisers and invariant priors. In this paper, we challenge this tradition through the alternative canonicalization perspective: first map each sample to an orbit representative with a canonical pose or order, train an unconstrained (non-equivariant) diffusion or flow model on the canonical slice, and finally recover the invariant distribution by sampling a random symmetry transform at generation time. Building on a formal quotient-space perspective, our work provides a comprehensive theory of canonical diffusion by proving: (i) the correctness, universality and superior expressivity of canonical generative models over invariant targets; (ii) canonicalization accelerates training by removing diffusion score complexity induced by group mixtures and reducing conditional variance in flow matching. We then show that aligned priors and optimal transport act complementarily with canonicalization and further improves training efficiency. We instantiate the framework for molecular graph generation under $S_n \times SE(3)$ symmetries. By leveraging geometric spectra-based canonicalization and mild positional encodings, canonical diffusion significantly outperforms equivariant baselines in 3D molecule generation tasks, with similar or even less computation. Moreover, with a novel architecture Canon, CanonFlow achieves state-of-the-art performance on the challenging GEOM-DRUG dataset, and the advantage remains large in few-step generation.
❌