GeneralResearch

GPT-4.1 and Llama 3.3 70 fail to detect clinically relevant errors in radiology reports in zero-shot evaluation.

European RadiologyJun 19

OBJECTIVES: To evaluate whether GPT-4.1 and Llama 3.3 70B, large language models (LLMs) assessed in zero-shot, baseline configurations, detect and categorize clinically consequential errors across types that range from pattern-based to reasoning-dependent. MATERIALS AND METHOD...

  • Source: European Radiology (PRIMARY)
  • Type: Research

Automated summary

RadPigeon summaries are original and for information only. They are not clinical advice.