❌

Normal view

Large Language Models’ Clinical Decision-Making on When to Perform a Kidney Biopsy: Comparative Study

Background: Artificial intelligence (AI) and Large Language models (LLMs) are increasing in sophistication and are being integrated into many disciplines. The potential for LLMs to augment clinical decisions is an evolving area of research. Objective: This study compared the responses of over 1000 kidney specialist physicians (nephrologists) to outputs of commonly used LLMs using a questionnaire determining when a kidney biopsy should be performed. Methods: This research group completed a large online questionnaire for nephrologists to determine when a kidney biopsy should be performed. The questionnaire was co-designed with patient participation, refined through multiple iterations, then piloted locally before international dissemination. It was the largest international study in the field and demonstrated variation between human clinicians in biopsy propensity relating to human factors such as sex and age, as well as systemic factors such as country, job seniority and technical proficiency. The same questions were put to both human doctors and LLMs in an identical order in a single session. Eight commonly used LLMs were interrogated: Chat GPT 3.5, Mistral Hugging Face, Perplexity, Microsoft Co-pilot, Llama 2, GPT 4.0, MedLM and Claude 3. The most common response given by clinicians (human mode) to each question was taken as the baseline for comparison. Questionnaire responses to the indications and contraindications for biopsy generated a score (0-44) reflecting biopsy propensity, in which a higher score was used as a surrogate marker for an increased tolerance of potential associated risks. Results: The ability of LLMs to reproduce human expert consensus varied widely with some models demonstrating a balanced approach to risk in a similar manner to humans, whilst other models reported outputs at either end of the spectrum for risk tolerance. In terms of agreement with the human mode, Chat GPT 3.5 and GPT 4.0 (Open AI) had the highest levels of alignment, with the human mode selected in 6/11 questions. The total biopsy propensity score generated from the human mode was 23/44. Both Open AI models produced similar propensity scores between 22 and 24, however Llama 2 and MS Co-pilot also reported scores within this range, but with poorer response alignment to the human mode at only 2/11 questions. The most risk averse model in this study was MedLM with a propensity score of 11 and the least risk averse model was Claude 3 with a score of 34. Conclusions: LLM outputs demonstrated a modest ability to replicate human clinical decision making in this study, however the performance varied widely between LLM models. Questions with more uniform human responses produced LLM outputs with greater alignment, whereas in questions with low levels of human consensus there was poor output alignment. This may limit the practical use of LLMs in real world clinical practice.

STAMP: Single-cell transcriptomics analysis and multimodal profiling through imaging

Single-cell transcriptomics analysis and multimodal profiling (STAMP) by imaging enables single-cell analysis of cells in suspension without the need for sequencing. The markedly reduced costs and flexible experimental designs support the profiling of millions of cells or the large-scale multiplexing of conditions, perturbations, and sample types.

Human interpretable grammar encodes multicellular systems biology models to democratize virtual cell laboratories

We developed a plain text modeling language—a cell behavior hypothesis grammar—to easily build virtual cell models and connect them to data, helping scientists to unlock the hidden dynamics of tissues. We provide examples showing how to use them in virtual experiments exploring how cancer responds to the cells in its environment and how the brain forms layers in development.

Google Gemini: Everything you need to know about the generative AI models

27 February 2025 at 10:09

Gemini is Google’s long-promised, next-gen generative AI model family.

© 2024 TechCrunch. All rights reserved. For personal use only.

  • ✇TechCrunch
  • OpenAI used this subreddit to test AI persuasion Maxwell Zeff
    OpenAI used the subreddit, r/ChangeMyView, to create a test for measuring the persuasive abilities of its AI reasoning models. The company revealed this in a system card — a document outlining how an AI system works — that was released along with its new “reasoning” model, o3-mini, on Friday. Millions of Reddit users are members […] © 2024 TechCrunch. All rights reserved. For personal use only.
     

OpenAI used this subreddit to test AI persuasion

1 February 2025 at 07:47

OpenAI used the subreddit, r/ChangeMyView, to create a test for measuring the persuasive abilities of its AI reasoning models. The company revealed this in a system card — a document outlining how an AI system works — that was released along with its new “reasoning” model, o3-mini, on Friday. Millions of Reddit users are members […]

© 2024 TechCrunch. All rights reserved. For personal use only.

What Trump 2.0 means for science: the likely winners and losers

Nature, Published online: 15 January 2025; doi:10.1038/d41586-025-00052-z

The incoming US president is expected to gut support for research on the environment and infectious diseases, but could buoy work in artificial intelligence, quantum research and space exploration.

Cancer biomarkers: Emerging trends and clinical implications for personalized treatment

Cancer biomarkers have transformed oncology, enabling treatments tailored to each tumor’s unique profile. This review highlights the field’s progress due to advancements in understanding cancer biology, testing methods, and understanding of the immune microenvironment to advance precision oncology for improved patient outcomes.

Translation of tissue-based artificial intelligence into clinical practice: from discovery to adoption

Oncogene, Published online: 24 October 2023; doi:10.1038/s41388-023-02857-6

Translation of tissue-based artificial intelligence into clinical practice: from discovery to adoption

Histone demethylase KDM5D upregulation drives sex differences in colon cancer

Nature, Published online: 21 June 2023; doi:10.1038/s41586-023-06254-7

A murine colorectal cancer (CRC) model shows that mutant KRAS-STAT4-mediated upregulation of Y chromosome KDM5D contributes to the sex differences in KRAS-mutant CRC, providing an actionable therapeutic strategy for metastasis risk reduction for men afflicted with KRAS-mutant CRC.

Spatial epigenome–transcriptome co-profiling of mammalian tissues

Nature, Published online: 15 March 2023; doi:10.1038/s41586-023-05795-1

The authors present two technologies for spatially resolved, genome-wide, joint profiling of the epigenome and transcriptome by cosequencing chromatin accessibility and gene expression, or histone modifications and gene expression on the same tissue section at near-single-cell resolution.
❌