BreastAI / InformaticsResearchTrainee
Reasoning LLMs surpass conventional models for BI-RADS educational question answering
Frontiers in medicine2w ago
Reasoning large language models (LLMs) scored higher than conventional ones for BI-RADS educational questions (median 3.7 vs 2.7, P<0.001). Performance fell for multifaceted clinical scenarios (reasoning models Δmedian -1.3). ChatGPT-o1 and Deepseek-R1 led. Language affected som…
- Reasoning LLMs outperformed conventional models with a median Likert score of 3.7 (IQR 2.7-4.0) versus 2.7 (2.0-3.7), P<0.001.
- Both model categories showed significantly lower scores on multifaceted clinical scenario questions; reasoning models dropped from median 4.0 to 2.7 (Δ -1.3, P<0.001).
- Question language (English/Chinese) did not affect reasoning models but impacted some conventional models (ChatGPT-3.5, Deepseek-V3, Gemini2.0-Flash).
Related reporting systems
Automated summary
RadPigeon summaries are original and for information only. They are not clinical advice.