GeneralAI / InformaticsResearch
Systematic Review: Wide Performance Variation with Small Clinical Language Models, Evidence Base Still Insufficient
Journal of the American Medical Informatics Association : JAMIA2d ago
SLMs (≤4B parameters) showed relative task scores from 0.30 to 2.36 across 9 studies, but the evidence is too heterogeneous to compare with larger models, and safety metrics like calibration remain unreported.
- Relative task scores for small language models ranged widely (0.30–2.36), precluding any pooled comparison with larger models.
- Hallucination was evaluated in only 6 of 11 studies, calibration error or epistemic uncertainty in none, and on-device inference timing in just 2.
- A 4B-parameter model requires an estimated 9.3 GB memory, suggesting single-GPU feasibility, but actual deployment was not demonstrated.
Automated summary
RadPigeon summaries are original and for information only. They are not clinical advice.