GeneralAI / InformaticsResearch

Chain-of-thought prompting improves AI detection of radiology report errors across seven LLMs

European radiology experimentalyesterday

RadCoT, radiologist-style chain-of-thought prompting, lifted mean error-detection F1 from 0.77 to 0.85 (p=0.003) across seven LLMs on 1,200 reports. GPT-4o hit F1 0.93; open-source Llama-3.3-70B reached 0.89, close to commercial standard prompting.

  • 1,200 reports (900 error-containing, 300 error-free) from a QA repository were evaluated across radiography, ultrasound, CT, and MRI.
  • Errors were clinician-validated and categorized into five types by experienced radiologists; interpretation errors gained most, with F1 improving from 0.57 to 0.75.
  • All seven LLMs improved with RadCoT; open-source Llama-3.3-70B with RadCoT matched GPT-4o standard prompting (Holm-adjusted p=0.28).

Automated summary

RadPigeon summaries are original and for information only. They are not clinical advice.