❌

Normal view

Comparative Evaluation of Generative AI Models for Chest Radiograph Report Generation in the Emergency Department

arXiv:2512.00271v1 Announce Type: cross Abstract: Purpose: To benchmark open-source or commercial medical image-specific VLMs against real-world radiologist-written reports. Methods: This retrospective study included adult patients who presented to the emergency department between January 2022 and April 2025 and underwent same-day CXR and CT for febrile or respiratory symptoms. Reports from five VLMs (AIRead, Lingshu, MAIRA-2, MedGemma, and MedVersa) and radiologist-written reports were randomly presented and blindly evaluated by three thoracic radiologists using four criteria: RADPEER, clinical acceptability, hallucination, and language clarity. Comparative performance was assessed using generalized linear mixed models, with radiologist-written reports treated as the reference. Finding-level analyses were also performed with CT as the reference. Results: A total of 478 patients (median age, 67 years [interquartile range, 50-78]; 282 men [59.0%]) were included. AIRead demonstrated the lowest RADPEER 3b rate (5.3% [76/1434] vs. radiologists 13.9% [200/1434]; P<.001 whereas other vlms showed higher disagreement rates p clinical acceptability was the highest with airead vs. radiologists while performed worse hallucinations were rare comparable to but frequent models language clarity lingshu and medversa compared sensitivity varied substantially across for common findings: maira-2 medgemma conclusion: medical cxr report generation exhibited variable performance in quality diagnostic measures.>
❌