Neuro / Head & NeckAI / InformaticsResearch

DeepSeek-R1 beats ChatGPT-o1 in neuroradiology academic writing quality but fabricates more citations

Digital health3d ago

DeepSeek-R1 (open-source) outperformed ChatGPT-o1 (proprietary) in academic writing quality for neuroradiology (mean Likert 3.23 vs 3.02, p=0.021). However, citation fabrication was higher (36.7% vs 2.6%). Both models require rigorous human review.

  • DeepSeek-R1 outperformed in reasoning depth (p=0.015), contextual coherence (p=0.043), subtlety (p=0.037), and evidence integration (p=0.027).
  • Citation reliability was poor; DeepSeek-R1 cited more real publications (55.1% vs 33.3%) but also had a higher confirmed fabrication rate (36.7% vs 2.6%, p<0.001).
  • Inter-rater agreement was poor for Likert-scale items (ICC=0.466) and substantial for binary items (κ=0.746).

Automated summary

RadPigeon summaries are original and for information only. They are not clinical advice.