New model, old risks: sociodemographic bias and adversarial hallucinations vulnerability in GPT-5
4 April 2026 at 08:00
npj Digital Medicine, Published online: 04 April 2026; doi:10.1038/s41746-026-02584-8
We re-evaluated GPT-5 using our published pipelines: 500 emergency vignettes across 32 sociodemographic labels for bias, and adversarial prompts with fabricated details. GPT-5 showed no measurable improvement over GPT-4o in sociodemographic-linked decision variation, with several LGBTQIA+ groups flagged for mental-health screening in 100% of cases. Adversarial hallucination rates were higher (65% vs 53% for GPT-4o); a mitigation prompt reduced this to 7.67%.