GeneralResearch
GPT-4.1 and Llama 3.3 70 fail to detect clinically relevant errors in radiology reports in zero-shot evaluation.
European Radiology1w ago
OBJECTIVES: To evaluate whether GPT-4.1 and Llama 3.3 70B, large language models (LLMs) assessed in zero-shot, baseline configurations, detect and categorize clinically consequential errors across types that range from pattern-based to reasoning-dependent. MATERIALS AND METHOD...
- Source: European Radiology (PRIMARY)
- Type: Research
RadPigeon summaries are original and for information only. They are not clinical advice.
