GeneralAI / InformaticsResearch
Large language models cannot reproduce radiology residency rank lists from applications alone but align well once interviews are added
Journal of the American College of Radiology : JACRyesterday
LLMs (large language models) failed to reproduce a radiology residency rank list from applications alone (median τ 0.15–0.36), but adding human interview scores raised agreement to τ 0.84–0.93. Application-only LLM rankings placed female, international, and non-MD applicants low…
- Single-institution pilot: 148 applicants for 7 positions; seven LLM configurations generated 140 rank lists under application-only and interview-score conditions.
- Sorting by summed interviewer score alone reproduced the final list at τ-b=0.83, whereas sorting by USMLE Step 2 CK score alone gave τ-b=0.12.
- The authors note the study does not establish whether the LLM or human committee produced the better list; AI should not be used as a primary rank-order list generator.
Automated summary
RadPigeon summaries are original and for information only. They are not clinical advice.