GeneralResearch

GPT-4.1 and Llama 3.3 70 fail to detect clinically relevant errors in radiology reports in zero-shot evaluation.

European Radiology1w ago

OBJECTIVES: To evaluate whether GPT-4.1 and Llama 3.3 70B, large language models (LLMs) assessed in zero-shot, baseline configurations, detect and categorize clinically consequential errors across types that range from pattern-based to reasoning-dependent. MATERIALS AND METHOD...

Read the source

RadPigeon summaries are original and for information only. They are not clinical advice.