Musculoskeletal (MSK)AI / InformaticsResearchTrainee

ChatGPT-5.0 outperforms GPT-4o but remains inferior to clinicians for Kellgren-Lawrence grading of knee OA on radiographs

Skeletal radiologyyesterday

For radiographic knee OA grading, ChatGPT-5.0 outperforms GPT-4o but remains inferior to clinicians. Binary detection sensitivity 0.96 but specificity limited. Not suitable for standalone assessment.

  • Agreement with the reference standard was highest for the radiologist (weighted κ = 0.87); both LLMs showed moderate agreement only.
  • Per-grade classification was most accurate for KL grades 0 and 4, but limited for grades 1–2.
  • Intraobserver repeatability: reference reader κ = 0.881, ChatGPT-5.0 κ = 0.591 vs GPT-4o κ = 0.485.

Automated summary

RadPigeon summaries are original and for information only. They are not clinical advice.