❌

Normal view

ShapE-GRPO: Shapley-Enhanced Reward Allocation for Multi-Candidate LLM Training

arXiv:2603.29871v1 Announce Type: new Abstract: In user-agent interaction scenarios such as recommendation, brainstorming, and code suggestion, Large Language Models (LLMs) often generate sets of candidate recommendations where the objective is to maximize the collective utility of the entire set rather than individual candidates independently. However, existing reinforcement learning post-training paradigms, such as Group Relative Policy Optimization (GRPO), typically assign the same set-level scalar reward to every candidate in the set. This leads to noisy training signals where poor candidates free-ride on the high reward produced by a single strong peer, resulting in suboptimal exploration. To address this, we propose Shapley-Enhanced GRPO (ShapE-GRPO). By leveraging the permutation-invariant nature of set-level utility, we derive a Shapley-enhanced formulation from cooperative game theory to decompose set-level rewards into granular, candidate-specific signals. We show that our formulation preserves the fundamental axioms of the Shapley value while remaining computationally efficient with polynomial-time complexity. Empirically, ShapE-GRPO consistently outperforms standard GRPO across diverse datasets with accelerated convergence during training.

Curcumol Induces G1 Phase Arrest in SK-Hep-1 Cells by Targeting SKP2-Mediated p27 Degradation

28 March 2026 at 18:00

Molecules. 2026 Mar 16;31(6):997. doi: 10.3390/molecules31060997.

ABSTRACT

CONTEXT: S-phase kinase-associated protein 2 (SKP2) is an oncogene and cell cycle regulator that mediates the ubiquitination of cell cycle regulators. Curcumol, a sesquiterpene natural product, has been reported to regulate SKP2-mediated ubiquitination degradation to overcome drug resistance in cancer cells. However, whether the cell cycle arrest effect of curcumol is related to SKP2's function in cancer cells and its mechanisms are still unclear.

OBJECTIVE: To investigate the role of SKP2 in curcumol-induced cell cycle arrest and its underlying mechanisms.

MATERIALS AND METHODS: Transcriptomic and proteomic analyses were used to screen the ubiquitination-related factors in curcumol treated hepatocellular carcinoma cells. Lentiviral overexpression, co-immunoprecipitation assays, ubiquitination analysis, and cell-line-derived xenograft (CDX) models were used to dissect the role and mechanisms of the identified ubiquitination-related factor in the cell cycle arrest effect of curcucmol.

RESULTS: Curcumol modulated the expression of CDK4, CDK6, Cyclin D1, p27 and SKP2. SKP2 was one candidate target of curcumol selected by multi-omics. Overexpressed SKP2 partially reversed curcumol-induced growth inhibition and G1-phase arrest. The increased expression of p27 induced by curcumol was attenuated by overexpressed SKP2. Curcumol impaired the interaction between SKP2 and p27, and led to the ubiquitination and degradation of p27. In vivo, curcumol effectively reduced tumor growth, and its antitumor effect was significantly mitigated by SKP2 overexpression.

DISCUSSION AND CONCLUSIONS: Curcumol reduced SKP2 expression, weakened the interaction between SKP2 and p27, inhibited degradation of p27, and then induced G1 phase cell-cycle arrest in SK-Hep-1 cells.

PMID:41900096 | PMC:PMC13029316 | DOI:10.3390/molecules31060997

Less is More: Improving LLM Alignment via Preference Data Selection

17 February 2026 at 13:00
arXiv:2502.14560v4 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) has emerged as a promising approach for aligning large language models with human preferences. While prior work mainly extends DPO from the aspect of the objective function, we instead improve DPO from the largely overlooked but critical aspect of data selection. Specifically, we address the issue of parameter shrinkage caused by noisy data by proposing a novel margin-maximization principle for dataset curation in DPO training. To further mitigate the noise in different reward models, we propose a Bayesian Aggregation approach that unifies multiple margin sources (external and implicit) into a single preference probability. Extensive experiments in diverse settings demonstrate the consistently high data efficiency of our approach. Remarkably, by using just 10\% of the Ultrafeedback dataset, our approach achieves 3\% to 8\% improvements across various Llama, Mistral, and Qwen models on the AlpacaEval2 benchmark. Furthermore, our approach seamlessly extends to iterative DPO, yielding a roughly 3\% improvement with 25\% online data, revealing the high redundancy in this presumed high-quality data construction manner. These results highlight the potential of data selection strategies for advancing preference optimization.
❌