CardiacAI / InformaticsResearch
DeepSeek and GPT-o1 Outperform GPT-4o in Patient Engagement for Cardiovascular Imaging Questions
JMIR formative researchyesterday
In 84 patient questions about cardiovascular imaging, DeepSeek and GPT-o1 gave answers with better user engagement and reassurance (96.4% and 98.8% rated good) than GPT-4o (53.6%), while accuracy, clarity, and completeness were similarly high across all three large language mode…
- Two cardiovascular radiologists scored 84 real patient questions across four domains using a 3-point rubric; accuracy, clarity, and completeness were high and comparable between models.
- Composite scores were lower for GPT-4o (median 11, IQR 10-12) than for DeepSeek and GPT-o1 (both median 12, IQR 11-12; P<.001).
- No unsafe statements were identified in any model, and no significant differences were found for accuracy (P=.91), clarity (P=.06), or completeness (P=.65).
Automated summary
RadPigeon summaries are original and for information only. They are not clinical advice.