Chest / ThoracicAI / InformaticsResearch

Retrieval-augmented generation boosts LLM-produced CTPA impression text similarity

PloS one2d ago

Dynamic retrieval of top-10 similar cases boosted GPT-4o’s ROUGE-1 F1 to 0.44-0.47 (vs 0.35-0.37 zero-shot) and LLaMA’s to 0.37-0.50 (vs 0.25-0.37) on 599 CTPA reports, all p<0.05.

  • In this retrospective study of 599 CTPA reports, dynamic retrieval of semantically matched examples significantly improved automated text-similarity metrics for LLM-generated impressions over zero-shot and fixed few-shot prompting.
  • The highest ROUGE-1 F1 scores were achieved with temperature 0 and k=10, but scores remained moderate (max 0.50), indicating room for improvement.
  • The authors note that radiologist verification remains necessary before clinical deployment.

Automated summary

RadPigeon summaries are original and for information only. They are not clinical advice.