❌

Normal view

Improving Safety Alignment via Balanced Direct Preference Optimization

arXiv:2603.22829v1 Announce Type: new Abstract: With the rapid development and widespread application of Large Language Models (LLMs), their potential safety risks have attracted widespread attention. Reinforcement Learning from Human Feedback (RLHF) has been adopted to enhance the safety performance of LLMs. As a simple and effective alternative to RLHF, Direct Preference Optimization (DPO) is widely used for safety alignment. However, safety alignment still suffers from severe overfitting, which limits its actual performance. This paper revisits the overfitting phenomenon from the perspective of the model's comprehension of the training data. We find that the Imbalanced Preference Comprehension phenomenon exists between responses in preference pairs, which compromises the model's safety performance. To address this, we propose Balanced Direct Preference Optimization (B-DPO), which adaptively modulates optimization strength between preferred and dispreferred responses based on mutual information. A series of experimental results show that B-DPO can enhance the safety capability while maintaining the competitive general capabilities of LLMs on various mainstream benchmarks compared to state-of-the-art methods. \color{red}{Warning: This paper contains examples of harmful texts, and reader discretion is recommended.

A multiomics Mendelian randomization study on PANoptosis-related genes and gastric cancer risk

J Int Med Res. 2026 Mar;54(3):3000605261430163. doi: 10.1177/03000605261430163. Epub 2026 Mar 16.

ABSTRACT

ObjectiveTo explore the potential involvement of PANoptosis-related genes in gastric cancer susceptibility through multiomics analyses.MethodsSummary-data-based Mendelian randomization was performed by integrating blood-derived methylation, gene expression, and protein quantitative trait loci data with genome-wide association study results. The findings were further evaluated in The Cancer Genome Atlas cohort, followed by protein-protein interaction analysis, drug prediction, and molecular docking.ResultsSummary-data-based Mendelian randomization and colocalization analyses identified several traits suggestively associated with gastric cancer risk. Genetically predicted higher expression of apoptosis and caspase activation inhibitor (AVEN) and hepatocyte growth factor (HGF) as well as higher HGF protein levels were associated with increased risk, whereas higher levels of protein phosphatase 2 regulatory subunit B beta (PPP2R2B) appeared to be protective. Multiomics integration suggested epigenetic regulation of HGF and PPP2R2B. The Cancer Genome Atlas analysis corroborated the dysregulation of these candidates, with high AVEN expression associated with poorer survival. Protein-protein interaction and drug prediction analyses highlighted functional networks and potential therapeutics, supported by molecular docking demonstrating strong HGF-binding affinities. However, these associations did not reach statistical significance in the independent validation cohort, possibly due to limited statistical power.ConclusionsThis study identified AVEN, HGF, and PPP2R2B as potential candidate genes for gastric cancer. These findings require further validation in larger cohorts.

PMID:41840829 | DOI:10.1177/03000605261430163

LR-SGS: Robust LiDAR-Reflectance-Guided Salient Gaussian Splatting for Self-Driving Scene Reconstruction

arXiv:2603.12647v1 Announce Type: cross Abstract: Recent 3D Gaussian Splatting (3DGS) methods have demonstrated the feasibility of self-driving scene reconstruction and novel view synthesis. However, most existing methods either rely solely on cameras or use LiDAR only for Gaussian initialization or depth supervision, while the rich scene information contained in point clouds, such as reflectance, and the complementarity between LiDAR and RGB have not been fully exploited, leading to degradation in challenging self-driving scenes, such as those with high ego-motion and complex lighting. To address these issues, we propose a robust and efficient LiDAR-reflectance-guided Salient Gaussian Splatting method (LR-SGS) for self-driving scenes, which introduces a structure-aware Salient Gaussian representation, initialized from geometric and reflectance feature points extracted from LiDAR and refined through a salient transform and improved density control to capture edge and planar structures. Furthermore, we calibrate LiDAR intensity into reflectance and attach it to each Gaussian as a lighting-invariant material channel, jointly aligned with RGB to enforce boundary consistency. Extensive experiments on the Waymo Open Dataset demonstrate that LR-SGS achieves superior reconstruction performance with fewer Gaussians and shorter training time. In particular, on Complex Lighting scenes, our method surpasses OmniRe by 1.18 dB PSNR.

AriadneMem: Threading the Maze of Lifelong Memory for LLM Agents

arXiv:2603.03290v1 Announce Type: cross Abstract: Long-horizon LLM agents require memory systems that remain accurate under fixed context budgets. However, existing systems struggle with two persistent challenges in long-term dialogue: (i) \textbf{disconnected evidence}, where multi-hop answers require linking facts distributed across time, and (ii) \textbf{state updates}, where evolving information (e.g., schedule changes) creates conflicts with older static logs. We propose AriadneMem, a structured memory system that addresses these failure modes via a decoupled two-phase pipeline. In the \textbf{offline construction phase}, AriadneMem employs \emph{entropy-aware gating} to filter noise and low-information message before LLM extraction and applies \emph{conflict-aware coarsening} to merge static duplicates while preserving state transitions as temporal edges. In the \textbf{online reasoning phase}, rather than relying on expensive iterative planning, AriadneMem executes \emph{algorithmic bridge discovery} to reconstruct missing logical paths between retrieved facts, followed by \emph{single-call topology-aware synthesis}. On LoCoMo experiments with GPT-4o, AriadneMem improves \textbf{Multi-Hop F1 by 15.2\%} and \textbf{Average F1 by 9.0\%} over strong baselines. Crucially, by offloading reasoning to the graph layer, AriadneMem reduces \textbf{total runtime by 77.8\%} using only \textbf{497} context tokens. The code is available at https://github.com/LLM-VLM-GSL/AriadneMem.
❌