BreastAI / InformaticsResearch

Gemini Scores Highest in Blinded Radiologist Evaluation of Arabic BI-RADS Translation

Diagnostics (Basel, Switzerland)3w ago

In a blinded evaluation, Gemini outperformed DeepSeek and ChatGPT for patient-friendly Arabic translation of BI-RADS breast imaging reports (mean rating 3.73 vs 3.54 vs 3.03, p<0.001). No laterality errors occurred in any model.

  • Radiologists rated Gemini (mean 3.73), DeepSeek (3.54), and ChatGPT (3.03) on eight domains; Gemini and DeepSeek each significantly outperformed ChatGPT across all domains (adjusted p<0.001).
  • No laterality errors were found among the 15 translations (five reports per model).
  • The authors caution that LLM translations should serve as radiologist-reviewed communication aids, not autonomous substitutes, to prevent diagnostic miscommunication.

Related reporting systems

Automated summary

RadPigeon summaries are original and for information only. They are not clinical advice.