Chest / ThoracicAI / InformaticsResearch
Large language model labels paired with CNNs screen chest X-rays for disease
Frontiers in digital health2w ago
GPT-4o achieved 92.9% accuracy in labeling chest X-ray reports as diseased vs no disease. Training ConvNeXt-Tiny with these labels yielded AUC 0.832 (95% CI 0.801-0.863) for automated screening on MIMIC-CXR, outperforming EfficientNet-B1 (p=0.014). LLM-derived weak supervision i…
- GPT-4o label quality against radiologist annotations: accuracy 92.9%; diseased class precision 97.4%, recall 90.5%; no-disease class precision 87.1%, recall 96.4%.
- ConvNeXt-Tiny achieved the highest AUC of 0.832 (95% CI [0.801, 0.863]) and balanced accuracy of 0.739, significantly higher than EfficientNet-B1 (AUC 0.797, p=0.014).
- The study was limited to binary disease/no-disease classification; further work is needed for label reliability and multi-label expansion.
Automated summary
RadPigeon summaries are original and for information only. They are not clinical advice.