❌

Normal view

PromptAudit: Auditing Prompt Sensitivity in LLM-Based Vulnerability Detection

arXiv:2605.24171v1 Announce Type: cross Abstract: Large language models are increasingly used for vulnerability detection, yet their reliability under different prompt formulations remains uncharacterized. We present PromptAudit, a controlled evaluation framework that isolates prompt effects by fixing the dataset, decoding, and parsing while varying only the prompting strategy. Using five prompting strategies across five open-weight models on 1,000 CVEs (6,074 code samples spanning 16 programming languages), we evaluate accuracy, recall, abstention, coverage, and effective F1. We find that standard chain-of-thought prompting achieves the strongest overall operational performance, while few-shot prompting provides model-dependent benefits that are most pronounced for prompt-sensitive models. In contrast, adaptive chain-of-thought frequently suppresses recall and self-consistency induces excessive abstention, sharply reducing effective performance. These results show that vulnerability detection behavior is jointly determined by the model and the prompt, and that prompt sensitivity is a first-class system property that must be explicitly characterized in evaluation and deployment.

Oncogenic and tumor-suppressive forces converge on a progenitor niche at the benign-to-malignant transition

Opposing oncogenic and tumor-suppressive forces establish a progenitor-like state that builds a self-reinforcing niche to drive benign-to-malignant transition in pancreatic cancer models. Disruption of p53 activity or KRAS inhibits malignancy by collapsing the niche.

A spatial atlas of the healthy human liver from live donors

Nature, Published online: 15 April 2026; doi:10.1038/s41586-026-10377-y

A human spatial atlas of gene expression in liver based on live donors shows marked porto–central zonation of hepatocytes and non-parenchymal cells, and transcriptomic changes in early steatosis.
❌