❌

Reading view

Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection

arXiv:2605.24834v1 Announce Type: cross Abstract: Large language model (LLM) safety classifiers such as Llama Guard are effective at detecting overtly harmful prompts but remain vulnerable to adversarial jailbreak attacks that disguise malicious intent through role-play scenarios, fictional framing, and indirect requests. We present Reflect-Guard, a method that augments LLM-based safety classifiers with chain-of-thought self-reflection capabilities through parameter-efficient fine-tuning. Our approach distills analytical reasoning from GPT-4o-mini into structured reflection annotations, then trains Llama-Guard-3-8B via QLoRA to generate logical self-reflections before issuing safety verdicts. Using only 1000 training examples and updating just 0.5% of model parameters (~42M), Reflect-Guard achieves substantial improvements on two challenging benchmarks. On WildGuardTest, F1 score improves from 0.770 to 0.842 (+7.2 pp), with recall on adversarial prompts increasing from 0.513 to 0.921 (+40.8 pp). On JailbreakBench, the attack success rate drops from 10.3% to 1.8%, representing an 82.5% relative reduction. These gains are especially pronounced on adversarial inputs, where the explicit reasoning step enables the model to see through obfuscation techniques that defeat standard pattern-matching approaches. Our results demonstrate that teaching safety classifiers to reason about adversarial intent, rather than simply classify surface patterns, is a promising direction for robust LLM safety.
  •  

MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents

arXiv:2602.02474v2 Announce Type: replace-cross Abstract: Most Large Language Model (LLM) agent memory systems rely on a small set of static, hand-designed operations for extracting memory. These fixed procedures hard-code human priors about what to store and how to revise memory, making them rigid under diverse interaction patterns and inefficient on long histories. To this end, we present \textbf{MemSkill}, which reframes these operations as learnable and evolvable memory skills, structured and reusable routines for extracting, consolidating, and pruning information from interaction traces. Inspired by the design philosophy of agent skills, MemSkill employs a \emph{controller} that learns to select a small set of relevant skills, paired with an LLM-based \emph{executor} that produces skill-guided memories. Beyond learning skill selection, MemSkill introduces a \emph{designer} that periodically reviews hard cases where selected skills yield incorrect or incomplete memories, and evolves the skill set by proposing refinements and new skills. Together, MemSkill forms a closed-loop procedure that improves both the skill-selection policy and the skill set itself. Experiments on LoCoMo, LongMemEval, HotpotQA, and ALFWorld demonstrate that MemSkill improves task performance over strong baselines and generalizes well across settings. Further analyses shed light on how skills evolve, offering insights toward more adaptive, self-evolving memory management for LLM agents.
  •  

Stromal ACTA2 Counteracts TCDD-Induced Hepatocarcinogenesis via Suppression of the PI3K-AKT-mTOR Pathway

J Hepatocell Carcinoma. 2026 May 10;13:586916. doi: 10.2147/JHC.S586916. eCollection 2026.

ABSTRACT

PURPOSE: 2,3,7,8-Tetrachlorodibenzo-p-dioxin (TCDD) is a persistent environmental pollutant that promotes hepatocellular carcinoma (HCC) through non-genotoxic mechanisms. However, stromal regulatory factors that counteract its tumor-promoting effects remain poorly defined. This study aimed to elucidate the role of actin alpha-2 (ACTA2) in TCDD-associated hepatocarcinogenesis.

METHODS: An integrative strategy combining network toxicology, Mendelian randomization, multi-omics and single-cell analyses, molecular docking and molecular dynamics simulations, along with in vitro experiments, was employed to investigate the functional role of ACTA2.

RESULTS: ACTA2 was identified as a stromal-associated factor linked to reduced HCC risk and improved patient survival. Single-cell and multi-omics analyses revealed that ACTA2 is predominantly expressed in hepatic stellate cells and fibroblast-like populations, reflecting tumor microenvironment composition rather than tumor cell-intrinsic expression. Functional enrichment analyses indicated that ACTA2 is associated with extracellular matrix remodeling and PI3K-AKT signaling. Molecular simulations demonstrated stable binding of TCDD to ACTA2 (Ξ”G_bind β‰ˆ -7.05 kcal/mol), suggesting potential structural perturbation. In vitro experiments showed that TCDD downregulated ACTA2 expression, promoted proliferation of LX-2 and cancer-associated fibroblasts (CAFs), and activated PI3K-AKT-mTOR signaling, whereas ACTA2 overexpression attenuated these effects.

CONCLUSION: ACTA2 acts as a context-dependent stromal regulator that modulates PI3K-AKT-mTOR signaling in TCDD-induced hepatocarcinogenesis. These findings highlight the importance of stromal remodeling in environmental carcinogenesis and suggest ACTA2 as a potential biomarker and therapeutic target in dioxin-associated HCC.

PMID:42148320 | PMC:PMC13175077 | DOI:10.2147/JHC.S586916

  •  
❌