Musculoskeletal (MSK)PediatricAI / InformaticsResearch
Vision-language models fall short in age estimation from dental radiographs
Clinical oral investigationsyesterday
Three VLMs (ChatGPT-4o, Copilot) achieved <35% agreement within ±1 year for chronological age from panoramic radiographs, with poorest performance at age extremes (7-8, 17-20 yrs). Cannot be recommended for clinical or forensic use.
- Overall agreement for all models was below 35% within ±1 year of true chronological age.
- Slightly better agreement was observed in individuals aged 9-10 years; no correct estimates for ages 7, 17-20 years with ChatGPT-4o.
Automated summary
RadPigeon summaries are original and for information only. They are not clinical advice.