BreastAI / InformaticsResearchTrainee

Reasoning LLMs surpass conventional models for BI-RADS educational question answering

Frontiers in medicine2w ago

Reasoning large language models (LLMs) scored higher than conventional ones for BI-RADS educational questions (median 3.7 vs 2.7, P<0.001). Performance fell for multifaceted clinical scenarios (reasoning models Δmedian -1.3). ChatGPT-o1 and Deepseek-R1 led. Language affected som…

  • Reasoning LLMs outperformed conventional models with a median Likert score of 3.7 (IQR 2.7-4.0) versus 2.7 (2.0-3.7), P<0.001.
  • Both model categories showed significantly lower scores on multifaceted clinical scenario questions; reasoning models dropped from median 4.0 to 2.7 (Δ -1.3, P<0.001).
  • Question language (English/Chinese) did not affect reasoning models but impacted some conventional models (ChatGPT-3.5, Deepseek-V3, Gemini2.0-Flash).

Related reporting systems

Automated summary

RadPigeon summaries are original and for information only. They are not clinical advice.