❌

Reading view

Large Language Models’ Clinical Decision-Making on When to Perform a Kidney Biopsy: Comparative Study

Background: Artificial intelligence (AI) and Large Language models (LLMs) are increasing in sophistication and are being integrated into many disciplines. The potential for LLMs to augment clinical decisions is an evolving area of research. Objective: This study compared the responses of over 1000 kidney specialist physicians (nephrologists) to outputs of commonly used LLMs using a questionnaire determining when a kidney biopsy should be performed. Methods: This research group completed a large online questionnaire for nephrologists to determine when a kidney biopsy should be performed. The questionnaire was co-designed with patient participation, refined through multiple iterations, then piloted locally before international dissemination. It was the largest international study in the field and demonstrated variation between human clinicians in biopsy propensity relating to human factors such as sex and age, as well as systemic factors such as country, job seniority and technical proficiency. The same questions were put to both human doctors and LLMs in an identical order in a single session. Eight commonly used LLMs were interrogated: Chat GPT 3.5, Mistral Hugging Face, Perplexity, Microsoft Co-pilot, Llama 2, GPT 4.0, MedLM and Claude 3. The most common response given by clinicians (human mode) to each question was taken as the baseline for comparison. Questionnaire responses to the indications and contraindications for biopsy generated a score (0-44) reflecting biopsy propensity, in which a higher score was used as a surrogate marker for an increased tolerance of potential associated risks. Results: The ability of LLMs to reproduce human expert consensus varied widely with some models demonstrating a balanced approach to risk in a similar manner to humans, whilst other models reported outputs at either end of the spectrum for risk tolerance. In terms of agreement with the human mode, Chat GPT 3.5 and GPT 4.0 (Open AI) had the highest levels of alignment, with the human mode selected in 6/11 questions. The total biopsy propensity score generated from the human mode was 23/44. Both Open AI models produced similar propensity scores between 22 and 24, however Llama 2 and MS Co-pilot also reported scores within this range, but with poorer response alignment to the human mode at only 2/11 questions. The most risk averse model in this study was MedLM with a propensity score of 11 and the least risk averse model was Claude 3 with a score of 34. Conclusions: LLM outputs demonstrated a modest ability to replicate human clinical decision making in this study, however the performance varied widely between LLM models. Questions with more uniform human responses produced LLM outputs with greater alignment, whereas in questions with low levels of human consensus there was poor output alignment. This may limit the practical use of LLMs in real world clinical practice.
  •  
❌