BreastAI / InformaticsResearch

Objective vs subjective evaluation of AI-generated mammogram reports reveals over- and underestimation risks

Japanese journal of radiologytoday

Objective and subjective evaluation of a mammogram-reporting AI found BI-RADS agreement 58.1%, mass 76.7%, calcification 81.4%, ROUGE-L F1 0.672, but some cases with high objective scores were rated low subjectively, flagged as over/underestimation.

  • An AI model (Qwen2.5 finetuned) generated mammogram reports; BI-RADS agreement was 58.1%, mass detection 76.7%, calcification 81.4%.
  • Objective metrics (ROUGE-L F1 0.672, BLEU 0.542) did not always correlate with subjective clinician ratings, revealing over- and underestimation in some samples.
  • The findings highlight that radiologists should not rely solely on automated metrics for report quality; multifaceted evaluation is essential.

Related reporting systems

Automated summary

RadPigeon summaries are original and for information only. They are not clinical advice.