GeneralAI / InformaticsResearch

Systematic Review: LLM Report Generation Shows Inconsistent Safety and Workflow Impact

Journal of medical Internet researchyesterday

Systematic review of 101 studies finds LLM-generated medical reports show moderate acceptance but persistent clinically significant errors and mixed workflow effects; evidence too heterogeneous to support autonomous use, requiring clinician supervision.

  • In a chest x-ray study, AI report acceptance was similar to radiologist reports (70.5% vs 73.3%), but false-negative findings were slightly higher (18.5% vs 17.8%).
  • In a clinician-collaboration chest x-ray study, AI reports were equivalent or preferred in 77.7% and 56.1% of cases across two datasets, yet clinically significant errors persisted.
  • Evidence remains too heterogeneous and biased to support meta-analysis; all studies had moderate to serious risk of bias, precluding claims of autonomous readiness.

Automated summary

RadPigeon summaries are original and for information only. They are not clinical advice.