❌

Reading view

CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents

arXiv:2605.25624v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven breakthroughs in domains such as math, tool-use, and software engineering, yet its extension to computer-use agents (CUAs) has been bottlenecked by the scarcity of scalable training data with deterministic rewards. Constructing such data for CUAs requires consistent task instruction, executable environment, and verifiable reward. However, hand-curated benchmarks achieve high reward fidelity but cover few applications and LLM-as-judge-based datasets scale broadly but lack reliable verification. We present CUA-Gym, a scalable pipeline that co-generates task instructions, environment states, and reward functions. Concretely, a Generator agent constructs the initial and golden environment states, and a separate Discriminator agent writes the reward function from the task specification. An orchestrator agent drives the two through iterative rounds upon execution. Generated tuples then pass a final filter combining LLM majority voting and agent rollouts, ensuring quality beyond the per-task adversarial loop. To address the scarcity of training environments, we further synthesize CUA-Gym-Hub, a broad suite of high-fidelity mock web applications grounded in real-world software-use distributions, expanding the scale of CUA RLVR data by magnitude. Using this pipeline, we construct CUA-Gym, a dataset of 32,112 verified RLVR training tuples grounded in 110 environments. Trained with GSPO on CUA-Gym, our CUA-Gym-A3B and CUA-Gym-A17B achieve 62.1% and 72.6% on OSWorld-Verified, outperforming prior open-source CUAs at comparable scales, with performance scaling smoothly in both data volume and environment diversity. The same checkpoints also improve on the held-out WebArena benchmark, indicating transfer beyond the training environments. We will open-source the full synthesis pipeline, dataset, CUA-Gym-Hub environments, and models.
  •  

From Prompt Optimization to Multi-Dimensional Credibility Evaluation: Enhancing Trustworthiness of Chinese LLM-Generated Liver MRI Reports -- with Preliminary Extension to Lung Cancer

arXiv:2510.23008v3 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated promising performance in generating diagnostic conclusions from imaging findings, thereby supporting radiology reporting, trainee education, and quality control. However, systematic guidance on how to optimize prompt design across different clinical contexts remains underexplored. Moreover, a comprehensive and standardized framework for assessing the trustworthiness of LLM-generated radiology reports is yet to be established. This study aims to enhance the trustworthiness of LLM-generated liver MRI reports by introducing a Multi-Dimensional Credibility Assessment (MDCA) framework and providing guidance on institution-specific prompt optimization. The proposed framework is applied to evaluate and compare the performance of several advanced LLMs, including Kimi-K2-Instruct-0905, Qwen3-235B-A22B-Instruct-2507, DeepSeek-V3, and ByteDance-Seed-OSS-36B-Instruct, using the SiliconFlow platform.
  •  

Kaempferol functionally reprograms CD47 signaling to promote cytoprotection and attenuate oxeiptosis in severe acute pancreatitis

Phytomedicine. 2026 May 15;157:158305. doi: 10.1016/j.phymed.2026.158305. Online ahead of print.

ABSTRACT

BACKGROUND: Severe acute pancreatitis (SAP) lacks targeted therapies, and massive loss of functional pancreatic acinar cells (PAC) drives mortality. Kaempferol (KA) possesses well-established anti-inflammatory and cytoprotective activities and is derived from herbal medicinal plants, but its direct molecular targets and mechanism of action in SAP remain undefined.

PURPOSE: To evaluate the protective effects of KA against SAP and to elucidate its molecular mechanism of specific action, with a focus on identifying the direct cellular target through which KA exerts its cytoprotective effects.

STUDY DESIGN: Gain‑/loss‑of‑function in vitro and PAC‑specific CD47 SAP mouse models, combined with multi‑omics screening and biophysical assays.

METHODS: CD47 manipulation (siRNA/overexpression) was performed in primary PACs and cell lines, combined with WT/CD47-/-/Mist1‑CD47‑iOE (PAC‑specific) mouse models. Network pharmacology, transcriptomics and proteomics were integrated to screen and validate KA's protective effects. Computational‑experimental approaches (molecular docking/dynamics, CETSA, SPR, co‑IP, pharmacological epistasis) characterized KA's allosteric modulation of CD47 signaling.

RESULTS: CD47 was upregulated in SAP; its knockout reduced PAC death via KEAP1/PGAM5/AIFM1-driven oxeiptosis. KA reduced PAC death across genotypes, afforded no extra benefit in CD47-KO, and was not overridden by CD47‑OE. Mechanistically, KA allosterically binds CD47 ectodomain, stabilizes the CD47‑ UBQLN1 complex, and redirects signaling from Gαi‑mediated death to Gβγ/ ERK/NRF2‑mediated survival. ERK inhibition attenuated KA's protection. KA's action was CD47‑dependent.

CONCLUSION: This study identifies anti-oxeiptosis as a novel pharmacological activity of KA in SAP. This is achieved through allosteric modulation of CD47, redirecting its signaling from death‑promoting to a protective axis via activating Gβγ/ERK/NRF2 to suppress oxeiptosis. These findings reveal the CD47‑oxeiptosis axis as a therapeutic target and position KA as a promising candidate for SAP therapy, adding a new mechanistic dimension to KA's known pharmacological profile.

PMID:42184499 | DOI:10.1016/j.phymed.2026.158305

  •  

Ferroptosis and macrophage polarization: mechanisms, interplay, and implications for medical applications

Cell Death Discovery, Published online: 23 May 2026; doi:10.1038/s41420-026-03147-2

Ferroptosis and macrophage polarization: mechanisms, interplay, and implications for medical applications
  •  
❌