❌

Reading view

Dynamic Dual-Granularity Skill Bank for Agentic RL

arXiv:2603.28716v2 Announce Type: replace Abstract: Agentic RL can benefit substantially from reusable experience, yet existing skill-based methods mainly extract trajectory-level guidance and often lack principled mechanisms for maintaining an evolving skill memory. We propose D2Skill, a dynamic dual-granularity skill bank for agentic RL that organizes reusable experience into task skills for high-level guidance and step skills for fine-grained decision support and error correction. D2Skill jointly trains the policy and skill bank through paired baseline and skill-injected rollouts under the same policy, using their performance gap to derive hindsight utility signals for both skill updating and policy optimization. Built entirely from training-time experience, the skill bank is continuously expanded through reflection and maintained with utility-aware retrieval and pruning. Experiments on ALFWorld, WebShop, and Search-Augmented QA tasks show that D2Skill substantially improves performance over skill-free baselines across models of different scales. Further ablations and analyses show that both dual-granularity skill modeling and dynamic skill maintenance are critical to these gains, while the learned skills exhibit higher utility, transfer across evaluation settings, and introduce only modest training overhead.
  •  

CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test

arXiv:2605.23491v2 Announce Type: replace-cross Abstract: Recently, Reinforcement Learning with Verifiable Rewards (RLVR) and Test-Time Scaling (TTS) have advanced LLM code generation through executable verification. Yet Ground-Truth Unit Tests (GT UTs) remain a bottleneck: SOTA RLVR methods require them for costly training, while existing TTS methods lose competitiveness without them. This motivates GT-free TTS, where existing methods directly use self-generated UTs to refine and select code candidates. Yet such UTs are often noisy or spuriously coupled with wrong code, and UT quality in turn cannot be validated without reliable code. The key challenge is therefore to jointly improve both. To this end, we present CoSPlay, a GT-free, training-free framework that jointly improves codes and UTs through cooperative self-play. It first explores diverse solution ideas and identifies their potential failure modes to produce discriminative UT ideas. It then uses bidirectional pass-count signals from the Code-UT execution matrix to iteratively prune or fix weak codes and refresh or replace unreliable UTs, letting the two pools co-evolve. Finally, when multiple codes remain tied at the highest pass count, it picks the final code from the largest output-consensus cluster, since correct codes agree on the same inputs while wrong codes diverge. Experiments on four challenging benchmarks show that CoSPlay on Qwen2.5-7B-Instruct improves average BoN from 22.1% to 33.2% and UT accuracy from 14.6% to 78.3%, matching or surpassing the RLVR model CURE-7B. When applied to CURE-7B, it further improves BoN by 5.7%. CoSPlay also generalizes across diverse backbones and outperforms GT-free TTS baselines under comparable token budgets, with continued gains as the budget scales up. These results suggest a scalable inference strategy for competitive code generation without any GT data.
  •  

Machine learning-based identification of key genes underlying sex differences in hepatocellular carcinoma and targeted drug screening

Biomed Rep. 2026 Apr 24;24(6):74. doi: 10.3892/br.2026.2147. eCollection 2026 Jun.

ABSTRACT

Hepatocellular carcinoma (HCC) shows a marked predominance in men, yet the molecular basis for this sex disparity remains unclear. The present study leveraged multi-omics data and machine learning algorithms to identify key genes associated with sex-specific differences in HCC and to screen for putative candidate compounds, aiming to provide new insights for sex-specific therapy. The mRNA expression data of male and female patients with HCC and paracancerous tissues were obtained from the GEO and TCGA databases. To mitigate overfitting, data were partitioned into independent training and testing sets. Candidate genes were screened by differential expression analysis and weighted gene co-expression network analysis. A total of four complementary algorithms, random forest, support vector machines, generalized linear models and extreme gradient boosting were used to identify key genes with high predictive capability. CYP17A1 and IRX3 were identified as the top differentially expressed core genes associated with HCC in men. Pan-cancer analysis showed that CYP17A1 was lowly expressed in the majority of tumors, but significantly highly expressed in HCC, rectal adenocarcinoma and gastric cancer (P<0.001). Functional cell-based assays showed that knockout of CYP17A1 inhibited the proliferation, migration and invasion ability of HCC cells (P<0.001). Immunohistochemistry showed that CYP17A1 protein expression was significantly increased in HCC tissues from male patients when compared with that in paracancerous tissues (P<0.001), whereas there was no significant difference in female patient tissues (P>0.05). Notably, while IRX3 was identified computationally, its functional role remains to be experimentally validated. Molecular docking predicted a potential interaction between the natural compound Saikosaponin A and the CYP17A1 protein, and cellular assays revealed that it dose-dependently inhibits HCC cell malignant phenotypes. The present study suggests that CYP17A1 is associated with sex differences in HCC, potentially via the androgen signaling axis. Furthermore, IRX3 emerges as a novel hypothesis-generating candidate gene. Finally, the findings of the present study highlight Saikosaponin A as a putative therapeutic candidate for male patients with HCC, warranting further target-dependency investigations.

PMID:42125766 | PMC:PMC13158723 | DOI:10.3892/br.2026.2147

  •  
❌