❌

Normal view

Mechanisms and reversal strategies of liver fibrosis: from regulation of cell fate to clinical translation

J Transl Med. 2026 May 25. doi: 10.1186/s12967-026-08312-w. Online ahead of print.

ABSTRACT

BACKGROUND: Liver fibrosis is a dynamic and reversible pathological process underlying chronic liver diseases, characterized by excessive extracellular matrix deposition and progressive hepatic architectural distortion. It acts as a critical precursor to cirrhosis, hepatic decompensation, and hepatocellular carcinoma, imposing a substantial global disease burden.

MAIN BODY: Accumulating evidence indicates that liver fibrosis is a highly plastic process governed by multicellular crosstalk, immune microenvironment remodeling, epigenetic-metabolic coupling, and mechanotransduction. This review outlines core cellular effectors and their heterogeneity revealed by single-cell omics, and highlights key regulatory layers including circadian rhythm, epigenetic imprinting, metabolic reprogramming, and the gut-liver axis, as well as etiology-specific differences in fibrosis progression, reversibility, and therapeutic response. We also summarize advances in non-invasive diagnosis and clinical translation of anti-fibrotic therapies, and discuss key bottlenecks leading to clinical trial failures.

CONCLUSION: A deeper understanding of cell fate regulation and multicellular ecosystem remodeling will facilitate the development of precise strategies to achieve meaningful fibrosis regression and improve long-term clinical outcomes in chronic liver diseases.

PMID:42185911 | DOI:10.1186/s12967-026-08312-w

When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs

arXiv:2605.24202v1 Announce Type: new Abstract: Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learning is unstable in ways that are poorly understood. We study when end-to-end RL training of multi-agent LLM workflows improves over their base models, comparing Shared-Policy training, where all roles update one policy, with Isolated-Policy training, where each role has its own parameters. Our experimental matrix spans Eval-Opt, Voting, and Orch-Workers workflows, math and code tasks, and three model scales (0.6B, 1.7B, 4B). We find that multi-agent RL usually improves over base models, but gains depend jointly on workflow, task, and scale, not on policy sharing alone. Isolated-Policy tends to reach higher peak accuracy yet more often falls off a terminal accuracy cliff, while Shared-Policy training does not eliminate failure; it redistributes failure into qualitatively different patterns. We then explain the strongest of these patterns through role-level gradient dynamics induced by workflow topology and policy routing: under Isolated-Policy, parallel same-role agents on shared prompts amplify per-role gradients and drive terminal degradation in Voting and Orch-Workers workflows; under Shared-Policy, asymmetric per-step gradient mass causes the shared policy to be captured by the dominant role, producing different failure signatures by task and workflow. Together, the empirical map and its underlying mechanisms show that policy sharing routes training pressure through different channels rather than offering uniform stability, making it a design choice with workflow- and task-conditional tradeoffs.

Don't Retrain, Just Reuse: Recovering Dual-Target Molecules from Single-Target Diffusion Models

arXiv:2605.25681v1 Announce Type: cross Abstract: Designing a single molecule that modulates two targets is a promising strategy for polypharmacology, but it remains substantially harder than standard single-target generation because one candidate must satisfy two binding requirements while preserving drug-likeness and synthesizability. Existing dual-target generative methods typically introduce dual-target capability by either retraining the generator or intervening in the diffusion process during sampling. The former can be costly and difficult to stabilize when dual-target supervision is sparse, while the latter may be sensitive to denoising-time target balancing and competing update directions. These limitations motivate a generator-preserving alternative that keeps the pretrained prior intact: can dual-target candidates instead be recovered from the input space of a frozen single-target diffusion model, without modifying its parameters or denoising dynamics? We formulate this task as a constrained multi-objective optimization problem and propose REUSE, a hierarchical evolutionary input-space search framework that combines pair-conditioned exploration with structured multi-stage selection to enforce dual-target affinity, chemical quality, and diversity. Experiments show that, compared with methods that modify the diffusion process, REUSE consistently improves dual-target affinity and balance, achieving a 20.9-percentage-point gain in Dual High Affinity over the strongest prior baseline while maintaining competitive molecular quality.

Design Conditions for Intra-Group Learning of Sequence-Level Rewards: Token Gradient Cancellation

arXiv:2604.13088v2 Announce Type: replace-cross Abstract: Reinforcement learning for multi-step reasoning with large language models (LLMs) typically relies on sparse terminal rewards, which creates a poorly conditioned credit-assignment problem: the final feedback is propagated uniformly across all intermediate decisions. This leads to high gradient variance, unstable training, and many ineffective updates, ultimately limiting sustained model improvement. We propose a counterfactual-comparison framework for credit assignment. For each input, the framework samples multiple reasoning trajectories and treats their differences as implicit approximations to alternative decisions. This yields an implicit process-level advantage estimator that converts sparse terminal rewards into step-sensitive learning signals. Building on this framework, we introduce Implicit Behavior Policy Optimization (IBPO), which substantially improves training stability and the performance ceiling on mathematical and code-reasoning benchmarks. Our results point to a promising direction for unlocking the reasoning potential of LLMs.

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction

arXiv:2604.17328v2 Announce Type: replace-cross Abstract: This paper investigates the length problem in sequence-level relative reinforcement learning. We observe that, although existing methods partially alleviate length-related phenomena, a more fundamental issue remains insufficiently characterized: the comparison units used during training lack inherent comparability. Building on this observation, we propose a new perspective: the length problem should not be viewed merely as a loss-scaling or normalization bias, but rather as a \emph{comparison unit construction} problem. We further establish a sample-construction-based training framework that, instead of applying post-hoc corrections to unequal-length responses, proactively constructs equal-length, alignable, and comparable training segments during generation. Within this framework, we propose EqLen, a concrete method applicable to group-relative comparison algorithms such as GRPO, GSPO, and RLOO. Through dual-track synchronous generation, prefix inheritance, and segment masking, EqLen efficiently collects effective equal-length training segments and enables stable

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning

arXiv:2605.05226v2 Announce Type: replace-cross Abstract: The central challenge of reinforcement learning for reasoning lies not only in the sparsity of outcome-level supervision, but more fundamentally in how to transform feedback provided only at the end of a sequence into fine-grained learning signals that can guide intermediate reasoning steps. Existing approaches either rely on outcome-level rewards for sequence-level optimization, which makes precise credit assignment difficult, or depend on externally constructed process supervision, which is costly and difficult to scale sustainably. To address this, we propose a new perspective: reinforcement learning for reasoning can be understood as the problem of internalizing outcome supervision into process supervision. From this perspective, we introduce a supervision-internalization method for reinforcement learning for reasoning, enabling the model to automatically extract process-level learning signals through identifying, correcting, and reusing failed reasoning trajectories, thereby achieving finer-grained policy optimization under outcome-only supervision. We further abstract this idea into a new training paradigm, in which the model continually generates and refines its own internal process supervision during reinforcement learning, opening a new path for fine-grained credit assignment in reinforcement learning for reasoning that differs from externally provided process supervision.

Reducing Credit Assignment Variance via Counterfactual Reasoning Paths

arXiv:2605.16302v2 Announce Type: replace-cross Abstract: Reinforcement learning for multi-step reasoning with large language models (LLMs) typically relies on sparse terminal rewards, which creates a poorly conditioned credit-assignment problem: the final feedback is propagated uniformly across all intermediate decisions. This leads to high gradient variance, unstable training, and many ineffective updates, ultimately limiting sustained model improvement. We propose a counterfactual-comparison framework for credit assignment. For each input, the framework samples multiple reasoning trajectories and treats their differences as implicit approximations to alternative decisions. This yields an implicit process-level advantage estimator that converts sparse terminal rewards into step-sensitive learning signals. Building on this framework, we introduce Implicit Behavior Policy Optimization (IBPO), which substantially improves training stability and the performance ceiling on mathematical and code-reasoning benchmarks. Our results point to a promising direction for unlocking the reasoning potential of LLMs.

GPNMB Drives Brain Metastasis by Sculpting a Pathological Endothelial-Immune Interactome

Cancer Discov. 2026 Apr 15. doi: 10.1158/2159-8290.CD-25-1663. Online ahead of print.

ABSTRACT

Brain metastases (BM) remain a devastating disease with dismal prognosis. How circulating tumor cells (CTCs) penetrate the blood brain barrier (BBB) and reprogram the brain microenvironment remain unclear. Using spatially resolved multi-omic profiling of CTCs and brain metastases, integrated with experimental and clinical analyses, we identified Glycoprotein Non-Metastatic Melanoma Protein B (GPNMB) as a CTC-secreted driver of vascular disruption and brain colonization. CBX3 upregulation induced GPNMB expression, which bound endothelial EGFR, triggering CBL-mediated ubiquitination and degradation. Attenuated EGFR signaling suppressed FTO and disrupted endothelial junctions via YTHDF2-dependent TJP1 m6A methylation. Remarkably, GPNMB-induced BBB remodeling promoted immune infiltration via CXCL12-CXCR4 axis, and induced time course-dependent T cell exhaustion within the brain microenvironment. Clinically, elevated CBX3⁺GPNMB⁺ CTCs and plasma CXCL12 were significantly associated with BM progression in lung cancer and melanoma. Therapeutically, dual blockade of GPNMB and PD1 enhanced anti-BM efficacy in mice, unveiling GPNMB as a promising target for precision immunotherapy.

PMID:41973996 | DOI:10.1158/2159-8290.CD-25-1663

❌