GeneralAI / InformaticsResearch

LLMs Boost Radiology Report Readability Yet Exhibit Output Instability

Academic radiologyyesterday

LLMs significantly improved radiology report readability (P<0.05), with DeepSeek-R1 performing best, but all models showed inherent output instability and information omission. Optimized structured prompts reduced model variance and improved translational accuracy, particularly…

  • All tested LLMs (DeepSeek-R1, ChatGPT-4.0, and a third unspecified model) exhibited inherent instability, omitted information, and generated risk-averse recommendations.
  • Optimized structured prompts substantially reduced output variance and improved accuracy, with the strongest effect in DeepSeek-R1 and ChatGPT-4.0.
  • Self-reported patient comprehension varied by age and education level, even though readability metrics improved.

Automated summary

RadPigeon summaries are original and for information only. They are not clinical advice.