❌

Reading view

Improving Safety Alignment via Balanced Direct Preference Optimization

arXiv:2603.22829v1 Announce Type: new Abstract: With the rapid development and widespread application of Large Language Models (LLMs), their potential safety risks have attracted widespread attention. Reinforcement Learning from Human Feedback (RLHF) has been adopted to enhance the safety performance of LLMs. As a simple and effective alternative to RLHF, Direct Preference Optimization (DPO) is widely used for safety alignment. However, safety alignment still suffers from severe overfitting, which limits its actual performance. This paper revisits the overfitting phenomenon from the perspective of the model's comprehension of the training data. We find that the Imbalanced Preference Comprehension phenomenon exists between responses in preference pairs, which compromises the model's safety performance. To address this, we propose Balanced Direct Preference Optimization (B-DPO), which adaptively modulates optimization strength between preferred and dispreferred responses based on mutual information. A series of experimental results show that B-DPO can enhance the safety capability while maintaining the competitive general capabilities of LLMs on various mainstream benchmarks compared to state-of-the-art methods. \color{red}{Warning: This paper contains examples of harmful texts, and reader discretion is recommended.
  •  

ESM1 drives cancer angiogenesis and bevacizumab resistance via trioleate synthesis

Neoplasia. 2026 May;75:101298. doi: 10.1016/j.neo.2026.101298. Epub 2026 Mar 20.

ABSTRACT

BACKGROUND: Hepatocellular carcinoma (HCC) exhibits high recurrence rates and limited therapeutic options. Endothelial cell-specific molecule 1 (ESM1) and angiopoietin-like 4 (ANGPTL4) are implicated in tumor progression, yet their synergistic role in HCC lipid metabolism and angiogenesis remains unexplored.

METHODS: We integrated multi-omics approaches, including RNA sequencing, metabolomics, and immunoprecipitation-mass spectrometry, in HCC cell lines and patient-derived xenograft models. Key experiments involved Co-IP, Western blotting, tube formation assays, and clinical tissue microarray analysis to validate the ESM1-ANGPTL4-FASN-trioleate axis.

RESULTS: ESM1 and ANGPTL4 formed a positive feedback loop, stabilizing fatty acid synthase (FASN) to promote trioleate synthesis. Trioleate activated the NF-ΞΊB/IL-17 pathway in HCC cells and upregulated CD99 in endothelial cells, driving angiogenesis. In vivo, ESM1/ANGPTL4 knockdown suppressed tumor growth, which was rescued by trioleate supplementation. Clinical data revealed elevated ESM1/ANGPTL4 expression in bevacizumab-resistant HCC, correlating with poor prognosis.

CONCLUSIONS: The ESM1-ANGPTL4-FASN-trioleate axis orchestrates metabolic reprogramming and endothelial activation, representing a promising therapeutic target. Future studies should explore combination therapies targeting this axis and overcoming bevacizumab resistance in HCC.

PMID:41864037 | PMC:PMC13019581 | DOI:10.1016/j.neo.2026.101298

  •  
❌