❌

Normal view

Frequently amplified SHANK2 promotes esophageal squamous cell carcinoma progression through the DVL/YAP axis

Oncogene, Published online: 31 August 2026; doi:10.1038/s41388-026-03974-8

Frequently amplified SHANK2 promotes esophageal squamous cell carcinoma progression through the DVL/YAP axis

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs

arXiv:2605.23965v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong performance on logical reasoning benchmarks, yet their reliability remains uncertain. Existing evaluations rely on static benchmarks, which fail to assess robustness under logically equivalent transformations and often overestimate reasoning capability. We propose LGMT (Logic-Grounded Metamorphic Testing), an oracle-free framework that leverages first-order logic (FOL) to evaluate LLM reasoning. By deriving metamorphic relations from formal logical equivalences, LGMT constructs semantically invariant test cases and detects reasoning defects through cross-case consistency checking. Experiments on six state-of-the-art LLMs show that LGMT exposes substantial hidden defects missed by traditional reference-based evaluations. We further find that models are particularly sensitive to symbol-level and conclusion-level variations, and that advanced prompting such as Few-shot CoT only partially mitigates these issues. These results suggest that LLM evaluation should move beyond isolated correctness toward robustness under logical invariance. LGMT provides a principled and scalable approach for diagnosing reasoning failures.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models

arXiv:2506.18543v2 Announce Type: replace-cross Abstract: The rapid proliferation of Large Language Models (LLMs) has heightened concerns regarding their exposure to jailbreak attacks, which craft adversarial inputs designed to elicit unsafe content. Although proprietary models such as GPT-4 have been extensively evaluated, the robustness of emerging open-source systems like DeepSeek remains insufficiently examined, despite their growing use in LLM applications. In this paper, we conduct the first comprehensive jailbreak analysis of the DeepSeek model family, comparing it with GPT-3.5 and GPT-4 through the HarmBench benchmark. We investigate seven representative attack methods across 510 harmful behaviors, organized along both functional and semantic dimensions. Findings indicate that DeepSeek provides partial resilience against optimization-driven attacks such as TAP-T, but also results in greater susceptibility to prompt-based and manually engineered adversarial inputs. In contrast, GPT-4 Turbo demonstrates more robust and consistent safety alignment across a wide range of behaviors, likely due to stronger safety optimization and reinforcement learning from human feedback. In addition, fine-grained behavioral analysis and case studies reveal that DeepSeek often fails to consistently apply safety constraints to adversarial prompts, leading to uneven refusal behaviors. Overall, our results highlight an inherent trade-off between model efficiency and alignment generalization, underscoring the importance of targeted safety tuning and robust alignment strategies to ensure secure deployment of open-source LLMs.

A Review of the Role of Zeqi Decoction in the Treatment of Non-Small Cell Lung Cancer

18 March 2026 at 18:00

J Multidiscip Healthc. 2026 Mar 11;19:584071. doi: 10.2147/JMDH.S584071. eCollection 2026.

ABSTRACT

Non-small cell lung cancer (NSCLC) is one of the malignant tumors with the highest incidence and mortality rates. Zeqi Decoction has the functions of "promoting diuresis and reducing swelling, resolving phlegm and dispersing nodules", embodying the unique approach of traditional Chinese medicine in treating lung cancer by "strengthening the body's resistance and eliminating pathogenic factors". Modern research shows that Zeqi Decoction exerts anti-NSCLC effects through multiple pathways and targets. In terms of the material basis of its efficacy, its active ingredients (such as diterpene esters and flavonoids contained in Zeqi) have the ability to directly inhibit the proliferation, invasion and migration of tumor cells and induce apoptosis. In terms of the mechanism of action, basic experiments have revealed that Zeqi Decoction can down-regulate the S100A9/STAT3 signaling pathway, inhibit the immunosuppressive activity of myelium-derived suppressor cells (MDSCs), reshape the tumor microenvironment, thereby enhancing the cytotoxic function of CD8⁺T cells, and can also regulate the EGFR/PI3K/Akt pathway to affect PD-L1 expression. Intervene in tumor immune escape; In terms of clinical transformation, the combination of Zexi Decoction with chemotherapy and targeted therapy can improve patients' symptoms such as cough and pleural effusion, prolong progression-free survival, and alleviate the toxic and side effects of Western medical treatment. In addition, Zexi Decoction also shows potential value in reversing drug resistance such as gemcitabine. At present, there are still problems such as the lack of standardized protocols and unclear molecular mechanisms in the research. In the future, it is necessary to combine new technologies such as network pharmacology and multi-omics analysis to deepen the research on the pharmacological material basis, dose-effect relationship and evidence-based medicine of Zeqi Decoction, so as to promote the clinical application and transformation of the combination of traditional Chinese and Western medicine in the treatment of NSCLC.

PMID:41847115 | PMC:PMC12991379 | DOI:10.2147/JMDH.S584071

ESPO: Entropy Importance Sampling Policy Optimization

arXiv:2512.00499v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a central component of post-training for large language models (LLMs), particularly for complex reasoning tasks that require stable optimization over long generation horizons. However, achieving performance at scale often introduces a fundamental trade-off between training stability and training efficiency. Token-level optimization applies fine-grained updates at the individual units, but is prone to high variance in gradient estimation, which can result in unstable training dynamics. In contrast, Sequence-level optimization often relies on aggressive clipping mechanisms to ensure stable updates. However, such design may discard a large fraction of valid training samples, leading to inefficient gradient utilization and reduced training efficiency. We refer to this phenomenon as gradient underutilization. In this work, we propose Entropy Importance Sampling Policy Optimization (ESPO), a novel framework that aims to combine fine-grained updates with stable training. ESPO decomposes sequences into groups based on predictive entropy, enabling (1) Entropy Grouping Importance Sampling to capture intra-sequence heterogeneity, and (2) Entropy Adaptive Clipping to dynamically allocate trust regions based on model uncertainty. Extensive experiments on mathematical reasoning benchmarks demonstrate that ESPO not only accelerates convergence but also achieves state-of-the-art performance, notably improving accuracy on the challenging mathematical benchmarks.
❌