BreastAI / InformaticsResearch
Objective vs subjective evaluation of AI-generated mammogram reports reveals over- and underestimation risks
Japanese journal of radiologytoday
Objective and subjective evaluation of a mammogram-reporting AI found BI-RADS agreement 58.1%, mass 76.7%, calcification 81.4%, ROUGE-L F1 0.672, but some cases with high objective scores were rated low subjectively, flagged as over/underestimation.
- An AI model (Qwen2.5 finetuned) generated mammogram reports; BI-RADS agreement was 58.1%, mass detection 76.7%, calcification 81.4%.
- Objective metrics (ROUGE-L F1 0.672, BLEU 0.542) did not always correlate with subjective clinician ratings, revealing over- and underestimation in some samples.
- The findings highlight that radiologists should not rely solely on automated metrics for report quality; multifaceted evaluation is essential.
Related reporting systems
Automated summary
RadPigeon summaries are original and for information only. They are not clinical advice.