❌

Normal view

FairMedQA: Benchmarking Bias in Large Language Models for Medical Question Answering

arXiv:2505.19562v2 Announce Type: replace Abstract: Large language models (LLMs) are approaching expert-level performance in medical question answering (QA), demonstrating strong potential to improve public healthcare. However, underlying biases related to sensitive attributes such as sex and race pose life-critical risks. The extent to which such sensitive attributes affect diagnosis remains an open question and requires comprehensive empirical investigation. Additionally, even the latest Counterfactual Patient Variations (CPV) benchmark can hardly distinguish the bias levels of different LLMs. To further explore these dynamics, we propose a new benchmark, FairMedQA, and benchmark 12 representative LLMs. FairMedQA contains 4,806 counterfactual question pairs constructed from 801 clinical vignettes. Our results reveal substantial accuracy disparity ranging from 3 to 19 percentage points across sensitive demographic groups. Notably, FairMedQA exposes biases that are at least 12 percentage points larger than those identified by the latest CPV benchmark, presenting superior benchmarking sensitivity. Our results underscore an urgent need for targeted debiasing techniques and more rigorous, identity-aware validation protocols before LLMs can be safely integrated into practical clinical decision-support systems.

Synonymous mutations promote tumorigenesis by disrupting m6A-dependent mRNA metabolism

13 February 2025 at 08:00
The impact of synonymous mutations remains elusive. Here, the authors demonstrate that synonymous mutations can promote tumorigenesis by disrupting post-transcriptional m6A modification. The findings provide fresh insights into understanding the genotype-phenotype relationship in cancer and beyond.

Interplay between gut microbial communities and metabolites modulates pan-cancer immunotherapy responses

Cell Metab. 2025 Jan 28:S1550-4131(24)00495-9. doi: 10.1016/j.cmet.2024.12.013. Online ahead of print.

ABSTRACT

Immune checkpoint blockade (ICB) therapy has revolutionized cancer treatment but remains effective in only a subset of patients. Emerging evidence suggests that the gut microbiome and its metabolites critically influence ICB efficacy. In this study, we performed a multi-omics analysis of fecal microbiomes and metabolomes from 165 patients undergoing anti-programmed cell death protein 1 (PD-1)/programmed death ligand 1 (PD-L1) therapy, identifying microbial and metabolic entities associated with treatment response. Integration of data from four public metagenomic datasets (n = 568) uncovered cross-cohort microbial and metabolic signatures, validated in an independent cohort (n = 138). An integrated predictive model incorporating these features demonstrated robust performance. Notably, we characterized five response-associated enterotypes, each linked to specific bacterial taxa and metabolites. Among these, the metabolite phenylacetylglutamine (PAGln) was negatively correlated with response and shown to attenuate anti-PD-1 efficacy in vivo. This study sheds light on the interplay among the gut microbiome, the gut metabolome, and immunotherapy response, identifying potential biomarkers to improve treatment outcomes.

PMID:39909032 | DOI:10.1016/j.cmet.2024.12.013

❌