Musculoskeletal (MSK)EmergencyAI / InformaticsResearch
Vision language models show unstable repeat fracture reads on forensic radiographs
Tomography (Ann Arbor, Mich.)2w ago
In 300 forensic long-bone radiographs, ChatGPT-5.2 achieved 83.0% accuracy (sensitivity 70.0%, specificity 96.0%), but all three vision language models showed unstable repeat reads one month later (within-model kappa 0.131-0.622). Reproducibility, not just accuracy, needs benchm…
- Emergency physicians had median sensitivity 92.0% vs forensic physicians 76.7%; forensic readers' specificity was tightly clustered (median 92.0%, range 87.3-100.0%).
- Claude Sonnet 4.5 had 46.3% accuracy and 11.3% specificity due to extreme false positives, while ChatGPT-5.2 and Gemini 3 Pro reached 83.0% and 77.7%.
- Structured fracture subtype descriptions were often correct once a fracture was detected, but end-to-end subtype accuracy remained low, and larger multicenter studies with independent external datasets are needed before forensic application.
Automated summary
RadPigeon summaries are original and for information only. They are not clinical advice.