❌

Reading view

When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs

arXiv:2605.24202v1 Announce Type: new Abstract: Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learning is unstable in ways that are poorly understood. We study when end-to-end RL training of multi-agent LLM workflows improves over their base models, comparing Shared-Policy training, where all roles update one policy, with Isolated-Policy training, where each role has its own parameters. Our experimental matrix spans Eval-Opt, Voting, and Orch-Workers workflows, math and code tasks, and three model scales (0.6B, 1.7B, 4B). We find that multi-agent RL usually improves over base models, but gains depend jointly on workflow, task, and scale, not on policy sharing alone. Isolated-Policy tends to reach higher peak accuracy yet more often falls off a terminal accuracy cliff, while Shared-Policy training does not eliminate failure; it redistributes failure into qualitatively different patterns. We then explain the strongest of these patterns through role-level gradient dynamics induced by workflow topology and policy routing: under Isolated-Policy, parallel same-role agents on shared prompts amplify per-role gradients and drive terminal degradation in Voting and Orch-Workers workflows; under Shared-Policy, asymmetric per-step gradient mass causes the shared policy to be captured by the dominant role, producing different failure signatures by task and workflow. Together, the empirical map and its underlying mechanisms show that policy sharing routes training pressure through different channels rather than offering uniform stability, making it a design choice with workflow- and task-conditional tradeoffs.
  •  

STREAM: A Data-Centric Framework for Mining High-Value Task-Oriented Dialogues from Streaming Media

arXiv:2605.25162v1 Announce Type: cross Abstract: Large language models for vertical domains are bottlenecked by the scarcity of complex, domain-specific task-oriented dialogues. Existing data acquisition pipelines face a persistent trilemma: expert annotation is expensive, real-world service conversations are constrained by privacy and commercial restrictions, and static corpora quickly become temporally stale. We propose Stream, a data-centric framework that leverages publicly available streaming media (live streams and short videos) to synthesize high-value service dialogues at scale. Stream mines authentic interaction signals from noisy streams and synthesizes conversations by integrating role-grounded persona construction with Conversational Blueprint construction; it further adopts retrieval-augmented generation (RAG) to support knowledge-aware responses. Based on Stream, we release StreamDial, a large-scale multi-domain dataset covering Automotive, Restaurant, and Hotel. StreamDial contains 87,498 dialogue sessions and 1,497,320 turns in total, with an average of 17.11 turns per session and a comparable scale across domains. Each session is organized as a structured quadruplet $\langle P_u, P_a, B, H \rangle$ that pairs dialogue history with explicit user/agent personas and a Conversational Blueprint, capturing realistic service behaviors such as requirement mining, constraint conflicts, negotiation, and recovery. Evaluations with automatic judges and downstream tasks show that StreamDial improves intrinsic dialogue quality over strong baselines, and models trained with StreamDial improve Dialogue State Tracking across backbones; we further report a completed human-evaluation set and encouraging multilingual transfer on Qwen3-8B under a controlled training budget. The data is released in https://github.com/hitxueliang/DialogDataSetBySTREAM.
  •  

AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration

arXiv:2605.20025v2 Announce Type: replace Abstract: Automating scientific discovery requires more than generating papers from ideas. Real research is iterative: hypotheses are challenged from multiple perspectives, experiments fail and inform the next attempt, and lessons accumulate across cycles. Existing autonomous research systems often model this process as a linear pipeline: they rely on single-agent reasoning, stop when execution fails, and do not carry experience across runs. We present AutoResearchClaw, a multi-agent autonomous research pipeline built on five mechanisms: structured multi-agent debate for hypothesis generation and result analysis, a self-healing executor with a \textsc{Pivot}/\textsc{Refine} decision loop that transforms failures into information, verifiable result reporting that prevents fabricated numbers and hallucinated citations, human-in-the-loop collaboration with seven intervention modes spanning full autonomy to step-by-step oversight, and cross-run evolution that converts past mistakes into future safeguards. On ARC-Bench, a 25-topic experiment-stage benchmark, AutoResearchClaw outperforms AI Scientist v2 by 54.7%. A human-in-the-loop ablation across seven intervention modes reveals that precise, targeted collaboration at high-leverage decision points consistently outperforms both full autonomy and exhaustive step-by-step oversight. We position AutoResearchClaw as a research amplifier that augments rather than replaces human scientific judgment. Code is available at https://github.com/aiming-lab/AutoResearchClaw.
  •  

BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting

arXiv:2605.17937v2 Announce Type: replace-cross Abstract: Quantitative backtesting is essential for evaluating trading strategies but remains hampered by high technical barriers and limited scalability. While Large Language Models (LLMs) offer a transformative path to automate this complex, interdisciplinary workflow through advanced code generation, tool usage, and agentic planning, the practical realization is significantly challenged by the current lack of a large-scale benchmark dedicated to automated quantitative backtesting, which hinders progress in this field. To bridge this critical gap, we introduce BacktestBench, the first large-scale benchmark for automated quantitative backtesting. Built from over 6 million real market records, it comprises 18,246 meticulously annotated question-answering pairs across four task categories: metrics calculation, ticker selection, strategy selection, and parameter confirmation. We also propose AutoBacktest, a robust multi-agent baseline that translates natural language strategies into reproducible backtests by coordinating a Summarizer for semantic factor extraction, a Retriever for validated SQL generation, and a Coder for Python backtesting implementation. Our evaluation on 23 mainstream LLMs, complemented by targeted ablations, identifies key factors that influence end-to-end performance and highlights the importance of grounded verification and standardized indicator representations.
  •  

A pathogen lncRNA secreted into rice sequesters a host miRNA for virulence

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10572-x

A fungal long non-coding RNA from Magnaporthe oryzae translocates into rice cells to sequester a host microRNA that normally represses PKR1, a negative immunity regulator, thereby facilitating infection and revealing a widespread RNA-based pathogen–host interaction mechanism.
  •  

Spatial transcriptomic-metabolic features of tumor foci and tumor capsule in microvascular invasion with hepatocellular carcinoma: A spatial multi-omics study

PLoS Med. 2026 May 15;23(5):e1004703. doi: 10.1371/journal.pmed.1004703. eCollection 2026 May.

ABSTRACT

BACKGROUND: Microvascular invasion (MVI) is closely related to the recurrence and metastasis of hepatocellular carcinoma (HCC), but the underlying cellular mechanism remains largely elusive. This study aims to elucidate the regional cellular discrepancy between MVI-positive (MVI+) and MVI-negative (MVI-) HCC by integrating Spatial transcriptomics (ST) and spatial metabolomics (SM).

METHODS AND FINDINGS: ST and SM were performed on six tissue samples from four patients (including 2 MVI+, 2 MVI-, and 2 paratumor tissues), with the integration of 79 public single-cell RNA sequencing datasets of HCC. Patient identity was used as a covariate in the linear equation for regional differentially expressed gene analysis with the ST data. Clinical validation was conducted through multiplex immunofluorescence staining in 79 patients, together with external validation in the cancer genome atlas (TCGA)-liver hepatocellular carcinoma (LIHC) cohort (n = 299) and an independent microarray dataset (n = 62). For cell-type-specific metabolic profiling, spatial transcriptomic-metabolic registration was performed. The functional roles of key metabolites were further validated in vitro using inflammatory cancer-associated fibroblasts (iCAFs) derived from hepatic stellate cells (HSCs) and primary CAFs through co-culture models and various functional assays assessing cell proliferation, migration, and invasion. In the tumor lesion, a malignant STMN1+HMGN2+GPC3+ cell subtype enriched in MVI+ HCC was identified, which exhibited enhanced proliferative activity and was associated with poor prognosis. This finding was further confirmed in a local cohort of 79 patients, where multiplex immunofluorescence staining for the three genes (STMN1, HMGN2, and GPC3) showed significantly higher expression in the MVI+ group than in the MVI- group (p = 0.046). Integrated SM analysis further revealed that this cell population underwent metabolic reprogramming characterized by suppressed glycerolipid metabolism. In the tumor capsule, iCAFs-related genes were downregulated in MVI+ cases, and iCAFs were located distally from the tumor boundary. Spatial metabolite mapping showed a strong correlation between taurine and iCAFs, and functional assays demonstrated that taurine promotes HCC proliferation and migration by suppressing iCAF activity. One limitation of this study is the small sample size of spatial omics data, which hinders a more complete molecular functional analysis of the STMN1+HMGN2+GPC3+ cell subtype and iCAFs in MVI+ HCC. Larger-scale ST cohorts are required to further validate and expand the findings of this study.

CONCLUSIONS: This integrative spatial atlas proposes a hypothesis that there exists a highly proliferative and metabolically reprogrammed malignant cell subtype in the tumor lesion of MVI+ HCC, and that taurine in the tumor capsule modulates iCAF activity to influence tumor progression. The exploratory results provide mechanistic insights into MVI-related HCC progression and offer potential avenues for targeted therapeutic intervention of MVI+ HCC.

PMID:42139279 | PMC:PMC13178920 | DOI:10.1371/journal.pmed.1004703

  •  

Integrated radiopathomics nomogram for predicting angiogenic microvascular patterns in NSCLC: a dual-center validation study

Ann Med. 2026 Dec;58(1):2654291. doi: 10.1080/07853890.2026.2654291. Epub 2026 Apr 17.

ABSTRACT

BACKGROUND: To develop and validate an integrated radiopathomics nomogram combining multiphase CT images, H&E-stained slides, and clinicopathological variables for predicting microvascular patterns (MVPs) in non-small cell lung cancer (NSCLC).

METHODS: We retrospectively included consecutive surgically resected NSCLC patients from two centers (n = 258). Patients from center 1 were randomly divided into training and internal validation cohorts, while patients from center 2 formed external validation cohort. CD34-immunohistochemistry was used as the reference standard for MVPs to classify patients into non-angiogenic alveolar (NAA) and non-NAA groups. Radiomics and pathomics features were extracted to construct single-phase radiomics, combined radiomics, and pathomics models. Rad-score and Path-score were derived from combined radiomics and pathomics models, respectively. Rad-score, Path-score, and clinicopathological independent predictors were integrated to develop a nomogram. Model performance was assessed by area under the curve (AUC), calibration curve, decision curve analysis (DCA), and DeLong test.

RESULTS: On multivariable analysis, histological grade was an independent predictor of NAA MVP. Combined radiomics model for predicting MVPs achieved AUCs of 0.863, 0.856, and 0.849 in training, internal validation, and external validation cohorts, showing better performance than single-phase models. Pathomics model yielded AUCs of 0.878, 0.860, and 0.833, however, its specificity markedly decreased in validation cohorts. Nomogram model achieved the superior performance across all cohorts, with AUCs of 0.911, 0.903, and 0.901, outperforming single-modality models (DeLong test: all p < 0.05).

CONCLUSION: The nomogram demonstrated high accuracy and robustness in predicting MVPs in NSCLC, offering a promising tool for characterizing the tumor microenvironment and supporting individualized treatment.

PMID:41992828 | DOI:10.1080/07853890.2026.2654291

  •  

Pan-neurodegeneration proteomics reveals disease subtypes and molecular signatures

A pan-neurodegeneration atlas built from multilayer, deep proteomics of 2,279 brain samples across 6 major diseases integrates whole proteome, detergent-insoluble proteome, and posttranslational modifications to enable intra- and inter-disease comparisons to reveal disease-specific subtypes and dysregulated pathways, while identifying shared changes such as GPNMB upregulation and NPTX2 downregulation.
  •  

Characterization and regulatory mechanism evaluation of C8orf33 in hepatocellular carcinoma through multiomics profiling

Discov Oncol. 2026 Apr 11. doi: 10.1007/s12672-026-04951-z. Online ahead of print.

ABSTRACT

BACKGROUND: Hepatocellular carcinoma (HCC) is a major cause of cancer-related mortality. Chromosome 8 open reading frame 33 (C8orf33) has been noted as a potential oncogenic factor in several cancers, but its biological roles and regulatory mechanism in HCC microenvironment remain unknown.

METHODS: We integrated bulk RNA sequencing, single-cell RNA sequencing (scRNA-seq), and spatial transcriptomics (ST) to characterize the expression landscape of C8orf33. We then performed C8orf33 loss-of-function studies in HCC cell lines, including in vitro phenotypic assays and subcutaneous xenografts.

RESULTS: C8orf33 was broadly overexpressed and associated with unfavorable prognosis across multiple Cancers. In HCC, higher C8orf33 aligned with advanced stage and shorter overall survival. C8orf33 knockdown reduced proliferation and migration, impaired tumorigenic capacity, and increased apoptosis. ScRNA-seq analyses identified a malignant population of Epi3 with high C8orf33 expression. Cell-cell communication analysis suggested that C8orf33-high Epi3 state was associated with an enriched MIF-CD74/CXCR4/CD44 signaling program toward macrophage populations with M2-like features. ST analyses further confirmed the colocalization of C8orf33 with malignant features in tumor cores. In Huh7 cells, C8orf33 knockdown was accompanied by reduced mRNA and protein levels of MIF and its receptor components. Consistently, xenografts derived from C8orf33-silenced cells showed lower expression of these MIF-axis components and reduced infiltration of CD163 and CD206-positive macrophages.

CONCLUSION: These results support a tumor-promoting association of C8orf33 in HCC and suggest a potential link to macrophage-associated immunomodulatory features, nominating C8orf33 as a candidate biomarker and therapeutic target.

PMID:41965457 | DOI:10.1007/s12672-026-04951-z

  •  

Characterization and regulatory mechanism evaluation of C8orf33 in hepatocellular carcinoma through multiomics profiling

Discov Oncol. 2026 Apr 11. doi: 10.1007/s12672-026-04951-z. Online ahead of print.

ABSTRACT

BACKGROUND: Hepatocellular carcinoma (HCC) is a major cause of cancer-related mortality. Chromosome 8 open reading frame 33 (C8orf33) has been noted as a potential oncogenic factor in several cancers, but its biological roles and regulatory mechanism in HCC microenvironment remain unknown.

METHODS: We integrated bulk RNA sequencing, single-cell RNA sequencing (scRNA-seq), and spatial transcriptomics (ST) to characterize the expression landscape of C8orf33. We then performed C8orf33 loss-of-function studies in HCC cell lines, including in vitro phenotypic assays and subcutaneous xenografts.

RESULTS: C8orf33 was broadly overexpressed and associated with unfavorable prognosis across multiple Cancers. In HCC, higher C8orf33 aligned with advanced stage and shorter overall survival. C8orf33 knockdown reduced proliferation and migration, impaired tumorigenic capacity, and increased apoptosis. ScRNA-seq analyses identified a malignant population of Epi3 with high C8orf33 expression. Cell-cell communication analysis suggested that C8orf33-high Epi3 state was associated with an enriched MIF-CD74/CXCR4/CD44 signaling program toward macrophage populations with M2-like features. ST analyses further confirmed the colocalization of C8orf33 with malignant features in tumor cores. In Huh7 cells, C8orf33 knockdown was accompanied by reduced mRNA and protein levels of MIF and its receptor components. Consistently, xenografts derived from C8orf33-silenced cells showed lower expression of these MIF-axis components and reduced infiltration of CD163 and CD206-positive macrophages.

CONCLUSION: These results support a tumor-promoting association of C8orf33 in HCC and suggest a potential link to macrophage-associated immunomodulatory features, nominating C8orf33 as a candidate biomarker and therapeutic target.

PMID:41965457 | DOI:10.1007/s12672-026-04951-z

  •  

Single-cell spatiotemporal dissection of the human maternal–fetal interface

Nature, Published online: 08 April 2026; doi:10.1038/s41586-026-10316-x

A single-cell multiomic atlas of the human maternal–fetal interface across pregnancy reveals cell types, states and spatial niches, developmental tissue architectures and transcriptional programmes, and identifies cell types with roles in pre-eclampsia, spontaneous preterm birth and miscarriage.
  •  

Superconductivity and electronic structures of nickelate thin film superstructures

Nature, Published online: 08 April 2026; doi:10.1038/s41586-026-10352-7

Engineered Ruddlesden–Popper nickelate superstructures show that specific Fermi surface features enable ambient-pressure superconductivity, linking structural configuration, electronic structure and superconducting behaviour. .
  •  

Harnessing foundation models for digital pathology without re-training

Nature Cancer, Published online: 03 April 2026; doi:10.1038/s43018-025-01108-9

Applications of digital pathology in clinical oncology have largely depended on the requirement for labeled data and model re-training. A study now presents PRET, a training-free framework with robust performance for pan-cancer diagnosis that adapts pathology foundation models to diverse tasks at inference stage, from screening and subtyping tasks to segmentation and metastasis detection tasks.
  •  

Beyond Matching to Tiles: Bridging Unaligned Aerial and Satellite Views for Vision-Only UAV Navigation

arXiv:2603.22153v2 Announce Type: replace-cross Abstract: Recent advances in cross-view geo-localization (CVGL) methods have shown strong potential for supporting unmanned aerial vehicle (UAV) navigation in GNSS-denied environments. However, existing work predominantly focuses on matching UAV views to onboard map tiles, which introduces an inherent trade-off between accuracy and storage overhead, and overlooks the importance of the UAV's heading during navigation. Moreover, the substantial discrepancies and varying overlaps in cross-view scenarios have been insufficiently considered, limiting their generalization to real-world scenarios. In this paper, we present Bearing-UAV, a purely vision-driven cross-view navigation method that jointly predicts UAV absolute location and heading from neighboring features, enabling accurate, lightweight, and robust navigation in the wild. Our method leverages global and local structural features and explicitly encodes relative spatial relationships, making it robust to cross-view variations, misalignment, and feature-sparse conditions. We also present Bearing-UAV-90k, a multi-city benchmark for evaluating cross-view localization and navigation. Extensive experiments show encouraging results that Bearing-UAV yields lower localization error than previous matching/retrieval paradigm across diverse terrains. Our code and dataset will be made publicly available.
  •  

Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills

arXiv:2512.16301v3 Announce Type: replace Abstract: Large language model (LLM) agents are moving beyond prompting alone. ChatGPT marked the rise of general-purpose LLM assistants, DeepSeek showed that on-policy reinforcement learning with verifiable rewards can improve reasoning and tool use, and OpenClaw highlights a newer direction in which agents accumulate persistent memory and reusable skills. Yet the research landscape remains fragmented across post-training, retrieval, memory, and skill systems. This survey studies these developments under a single notion of \emph{adaptation}: improving an agent, its tools, or their interaction after pretraining. We organize the field with a four-paradigm framework spanning agent adaptation and tool adaptation. On the agent side, A1 (tool-execution-signaled) and A2 (agent-output-signaled) improve the agent itself through supervised fine-tuning, preference optimization, and reinforcement learning with verifiable rewards. On the tool side, T1 (agent-agnostic) provides reusable pre-trained modules any agent can call, while T2 (agent-supervised) uses the agent's outputs to train memory systems, skill libraries, or lightweight subagents. Using this framework, we review post-training methods, adaptive memory architectures, and agent skills; compare their trade-offs in cost, flexibility, and generalization; and summarize evaluation practices across deep research, software development, computer use, and drug discovery. We conclude by outlining open problems in agent-tool co-adaptation, continual learning, safety, and efficient deployment.
  •  

Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation

arXiv:2602.13810v2 Announce Type: replace-cross Abstract: Learning expressive and efficient policy functions is a promising direction in reinforcement learning (RL). While flow-based policies have recently proven effective in modeling complex action distributions with a fast deterministic sampling process, they still face a trade-off between expressiveness and computational burden, which is typically controlled by the number of flow steps. In this work, we propose mean velocity policy (MVP), a new generative policy function that models the mean velocity field to achieve the fastest one-step action generation. To ensure its high expressiveness, an instantaneous velocity constraint (IVC) is introduced on the mean velocity field during training. We theoretically prove that this design explicitly serves as a crucial boundary condition, thereby improving learning accuracy and enhancing policy expressiveness. Empirically, our MVP achieves state-of-the-art success rates across several challenging robotic manipulation tasks from Robomimic and OGBench. It also delivers substantial improvements in training and inference speed over existing flow-based policy baselines.
  •  

Condition-Gated Reasoning for Context-Dependent Biomedical Question Answering

arXiv:2602.17911v2 Announce Type: replace-cross Abstract: Current biomedical question answering (QA) systems often assume that medical knowledge applies uniformly, yet real-world clinical reasoning is inherently conditional: nearly every decision depends on patient-specific factors such as comorbidities and contraindications. Existing benchmarks do not evaluate such conditional reasoning, and retrieval-augmented or graph-based methods lack explicit mechanisms to ensure that retrieved knowledge is applicable to given context. To address this gap, we propose CondMedQA, the first benchmark for conditional biomedical QA, consisting of multi-hop questions whose answers vary with patient conditions. Furthermore, we propose Condition-Gated Reasoning (CGR), a novel framework that constructs condition-aware knowledge graphs and selectively activates or prunes reasoning paths based on query conditions. Our findings show that CGR more reliably selects condition-appropriate answers while matching or exceeding state-of-the-art performance on biomedical QA benchmarks, highlighting the importance of explicitly modeling conditionality for robust medical reasoning.
  •  

GreenPhase: A Green Learning Approach for Earthquake Phase Picking

arXiv:2603.03344v1 Announce Type: cross Abstract: Earthquake detection and seismic phase picking are fundamental yet challenging tasks in seismology due to low signal-to-noise ratios, waveform variability, and overlapping events. Recent deep-learning models achieve strong results but rely on large datasets and heavy backpropagation training, raising concerns over efficiency, interpretability, and sustainability. We propose GreenPhase, a multi-resolution, feed-forward, and mathematically interpretable model based on the Green Learning framework. GreenPhase comprises three resolution levels, each integrating unsupervised representation learning, supervised feature learning, and decision learning. Its feed-forward design eliminates backpropagation, enabling independent module optimization with stable training and clear interpretability. Predictions are refined from coarse to fine resolutions while computation is restricted to candidate regions. On the Stanford Earthquake Dataset (STEAD), GreenPhase achieves excellent performance with F1 scores of 1.0 for detection, 0.98 for P-wave picking, and 0.96 for S-wave picking. This is accomplished while reducing the computational cost (FLOPs) for inference by approximately 83% compared to state-of-the-art models. These results demonstrate that the proposed model provides an efficient, interpretable, and sustainable alternative for large-scale seismic monitoring.
  •  

Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model

arXiv:2510.18573v2 Announce Type: replace-cross Abstract: We present Kaleido, a subject-to-video~(S2V) generation framework, which aims to synthesize subject-consistent videos conditioned on multiple reference images of target subjects. Despite recent progress in S2V generation models, existing approaches remain inadequate at maintaining multi-subject consistency and at handling background disentanglement, often resulting in lower reference fidelity and semantic drift under multi-image conditioning. These shortcomings can be attributed to several factors. Primarily, the training dataset suffers from a lack of diversity and high-quality samples, as well as cross-paired data, i.e., paired samples whose components originate from different instances. In addition, the current mechanism for integrating multiple reference images is suboptimal, potentially resulting in the confusion of multiple subjects. To overcome these limitations, we propose a dedicated data construction pipeline, incorporating low-quality sample filtering and diverse data synthesis, to produce consistency-preserving training data. Moreover, we introduce Reference Rotary Positional Encoding (R-RoPE) to process reference images, enabling stable and precise multi-image integration. Extensive experiments across numerous benchmarks demonstrate that Kaleido significantly outperforms previous methods in consistency, fidelity, and generalization, marking an advance in S2V generation.
  •  
❌