Musculoskeletal (MSK)AI / InformaticsResearch

ChatGPT-5 outperforms ChatGPT-4o in detecting lumbar spondylolisthesis on standing lateral radiographs but remains limited compared with spine surgeons

European journal of orthopaedic surgery & traumatology : orthopedie traumatologieyesterday

ChatGPT-5 outperformed ChatGPT-4o for detecting lumbar spondylolisthesis on standing lateral radiographs, with sensitivity 67.5% vs 49.4% and accuracy 61.5% vs 55.5%, but both remained limited compared with fellowship-trained spine surgeons.

  • In 200 standing lateral lumbar radiographs from VinDr-SpineXR, expert spine surgeon consensus reclassified 81% of pre-labeled positives as positive and 2% of pre-labeled negatives as positive.
  • Specificity was similar (57.3% for ChatGPT-5 vs 59.8% for ChatGPT-4o), and inter-rater agreement was higher for ChatGPT-5 (kappa 0.238 vs 0.091).
  • This was a multi-center pilot using a public imaging dataset; both LLMs underperformed fellowship-trained spine surgeons.

Automated summary

RadPigeon summaries are original and for information only. They are not clinical advice.