Musculoskeletal (MSK)PediatricAI / InformaticsResearch

Vision-language models fall short in age estimation from dental radiographs

Clinical oral investigationsyesterday

Three VLMs (ChatGPT-4o, Copilot) achieved <35% agreement within ±1 year for chronological age from panoramic radiographs, with poorest performance at age extremes (7-8, 17-20 yrs). Cannot be recommended for clinical or forensic use.

  • Overall agreement for all models was below 35% within ±1 year of true chronological age.
  • Slightly better agreement was observed in individuals aged 9-10 years; no correct estimates for ages 7, 17-20 years with ChatGPT-4o.

Automated summary

RadPigeon summaries are original and for information only. They are not clinical advice.