GeneralAI / InformaticsResearch
Systematic Review: LLM Report Generation Shows Inconsistent Safety and Workflow Impact
Journal of medical Internet researchyesterday
Systematic review of 101 studies finds LLM-generated medical reports show moderate acceptance but persistent clinically significant errors and mixed workflow effects; evidence too heterogeneous to support autonomous use, requiring clinician supervision.
- In a chest x-ray study, AI report acceptance was similar to radiologist reports (70.5% vs 73.3%), but false-negative findings were slightly higher (18.5% vs 17.8%).
- In a clinician-collaboration chest x-ray study, AI reports were equivalent or preferred in 77.7% and 56.1% of cases across two datasets, yet clinically significant errors persisted.
- Evidence remains too heterogeneous and biased to support meta-analysis; all studies had moderate to serious risk of bias, precluding claims of autonomous readiness.
Automated summary
RadPigeon summaries are original and for information only. They are not clinical advice.