❌

Reading view

When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs

arXiv:2605.24202v1 Announce Type: new Abstract: Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learning is unstable in ways that are poorly understood. We study when end-to-end RL training of multi-agent LLM workflows improves over their base models, comparing Shared-Policy training, where all roles update one policy, with Isolated-Policy training, where each role has its own parameters. Our experimental matrix spans Eval-Opt, Voting, and Orch-Workers workflows, math and code tasks, and three model scales (0.6B, 1.7B, 4B). We find that multi-agent RL usually improves over base models, but gains depend jointly on workflow, task, and scale, not on policy sharing alone. Isolated-Policy tends to reach higher peak accuracy yet more often falls off a terminal accuracy cliff, while Shared-Policy training does not eliminate failure; it redistributes failure into qualitatively different patterns. We then explain the strongest of these patterns through role-level gradient dynamics induced by workflow topology and policy routing: under Isolated-Policy, parallel same-role agents on shared prompts amplify per-role gradients and drive terminal degradation in Voting and Orch-Workers workflows; under Shared-Policy, asymmetric per-step gradient mass causes the shared policy to be captured by the dominant role, producing different failure signatures by task and workflow. Together, the empirical map and its underlying mechanisms show that policy sharing routes training pressure through different channels rather than offering uniform stability, making it a design choice with workflow- and task-conditional tradeoffs.
  •  

A pathogen lncRNA secreted into rice sequesters a host miRNA for virulence

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10572-x

A fungal long non-coding RNA from Magnaporthe oryzae translocates into rice cells to sequester a host microRNA that normally represses PKR1, a negative immunity regulator, thereby facilitating infection and revealing a widespread RNA-based pathogen–host interaction mechanism.
  •  

Integrated radiopathomics nomogram for predicting angiogenic microvascular patterns in NSCLC: a dual-center validation study

Ann Med. 2026 Dec;58(1):2654291. doi: 10.1080/07853890.2026.2654291. Epub 2026 Apr 17.

ABSTRACT

BACKGROUND: To develop and validate an integrated radiopathomics nomogram combining multiphase CT images, H&E-stained slides, and clinicopathological variables for predicting microvascular patterns (MVPs) in non-small cell lung cancer (NSCLC).

METHODS: We retrospectively included consecutive surgically resected NSCLC patients from two centers (n = 258). Patients from center 1 were randomly divided into training and internal validation cohorts, while patients from center 2 formed external validation cohort. CD34-immunohistochemistry was used as the reference standard for MVPs to classify patients into non-angiogenic alveolar (NAA) and non-NAA groups. Radiomics and pathomics features were extracted to construct single-phase radiomics, combined radiomics, and pathomics models. Rad-score and Path-score were derived from combined radiomics and pathomics models, respectively. Rad-score, Path-score, and clinicopathological independent predictors were integrated to develop a nomogram. Model performance was assessed by area under the curve (AUC), calibration curve, decision curve analysis (DCA), and DeLong test.

RESULTS: On multivariable analysis, histological grade was an independent predictor of NAA MVP. Combined radiomics model for predicting MVPs achieved AUCs of 0.863, 0.856, and 0.849 in training, internal validation, and external validation cohorts, showing better performance than single-phase models. Pathomics model yielded AUCs of 0.878, 0.860, and 0.833, however, its specificity markedly decreased in validation cohorts. Nomogram model achieved the superior performance across all cohorts, with AUCs of 0.911, 0.903, and 0.901, outperforming single-modality models (DeLong test: all p < 0.05).

CONCLUSION: The nomogram demonstrated high accuracy and robustness in predicting MVPs in NSCLC, offering a promising tool for characterizing the tumor microenvironment and supporting individualized treatment.

PMID:41992828 | DOI:10.1080/07853890.2026.2654291

  •  
❌