BreastAI / InformaticsResearch
Gemini Scores Highest in Blinded Radiologist Evaluation of Arabic BI-RADS Translation
Diagnostics (Basel, Switzerland)3w ago
In a blinded evaluation, Gemini outperformed DeepSeek and ChatGPT for patient-friendly Arabic translation of BI-RADS breast imaging reports (mean rating 3.73 vs 3.54 vs 3.03, p<0.001). No laterality errors occurred in any model.
- Radiologists rated Gemini (mean 3.73), DeepSeek (3.54), and ChatGPT (3.03) on eight domains; Gemini and DeepSeek each significantly outperformed ChatGPT across all domains (adjusted p<0.001).
- No laterality errors were found among the 15 translations (five reports per model).
- The authors caution that LLM translations should serve as radiologist-reviewed communication aids, not autonomous substitutes, to prevent diagnostic miscommunication.
Related reporting systems
Automated summary
RadPigeon summaries are original and for information only. They are not clinical advice.