❌

Normal view

Extracting Clinical Guideline Information Using Two Large Language Models: Evaluation Study

Background: The effective implementation of personalized pharmacogenomics (PGx) requires the integration of released clinical guidelines into decision support systems (CDSS) to facilitate clinical applications. Large language models (LLMs) can be valuable tools for automating information extraction and updates. Objective: To assess the effectiveness of repeated cross-comparisons and an agreement-threshold strategy in two advanced LLMs as supportive tools for updating information. Methods: The study evaluated the performance of two LLMs, GPT-4o and Gemini-1.5-Pro, in extracting PGx clinical guidelines and comparing their outputs with expert-annotated evaluations. The two LLMs classified 385 PGx clinical guidelines, with each recommendation tested 20 times per model. Accuracy was assessed by comparing the results with manually labeled data. Two prospectively defined strategies were employed to identify inconsistent predictions. The first involved repeated cross-comparison, flagging discrepancies between the most frequent classifications from each model. The second employed a consistency threshold strategy, which designated predictions appearing in less than 60% of the 40 combined outputs as unstable. Cases flagged by either strategy were subjected to manual review. This study also estimated the overall cost of model usage and was conducted between October 1 and November 30, 2024. Results: GPT-4o and Gemini-1.5-Pro yielded reproducibility rates of 97.8% (7,534/7,700) and 98.9% (7,612/7,700), respectively, based on the most frequent classification for each query. Compared with expert labels, GPT-4o achieved 93.5% accuracy (Cohen’s Kappa=0.90; P<.001 and gemini-1.5-pro accuracy kappa="0.89;" p both models demonstrated high overall performance with comparable weighted average f1 scores gemini: the generated consistent predictions for of guideline items reducing need manual review by among these agreed-upon cases only one diverged from expert labels. applying a predefined agreement-threshold strategy further reduced number priority to although error rate slightly increased inconsistencies identified through methods prompted prioritization minimize errors enhance clinical applicability. total combined cost using llms was conclusions: findings suggest that two can effectively streamline pgx integration into cdss while maintaining minimal cost. selective remains necessary this approach offers practical scalable solution classification in workflows.>
❌