Journal
Diagnostic Performance of Deep Learning in Medical Imaging: An Analytical Review and Quantitative Synthesis of the Evidence
المستخلص
Deep learning has produced a large body of studies reporting expert-level accuracy in the diagnostic interpretation of medical images, yet the headline numbers are easily misread. This analytical review examines the quantitative evidence with the tools of diagnostic-test methodology, asking not only how accurate deep-learning models are but what their reported accuracy means and how far it can be trusted. We define and apply the core metrics, sensitivity, specificity, the area under the receiver operating characteristic curve (AUC), Youden's index, likelihood ratios, and the diagnostic odds ratio, and we assemble reported pooled estimates from the principal meta-analyses across imaging domains. The best available synthesis of externally validated studies places the pooled sensitivity of deep-learning models at 87.0% and specificity at 92.5%, statistically indistinguishable from the 86.4% and 90.5% of health-care professionals assessed on the same samples, while domain-specific meta-analyses report AUCs from roughly 0.86 for respiratory imaging to above 0.93 for retinal disease. Behind these figures, however, lies a consistent methodological problem that our analysis foregrounds: extreme between-study heterogeneity, a scarcity of external validation and prospective testing, small and non-representative human comparators, poor adherence to reporting standards, and susceptibility to dataset shift and confounding, all of which tend to inflate apparent performance. We argue that the central analytical lesson is a gap between reported diagnostic accuracy and demonstrated clinical validity, and we set out the metrics, study designs, and reporting practices, including AI-specific reporting guidelines and prospective external validation, needed to close it. The review is quantitative and methodological in emphasis and is weighted toward classification tasks in radiology, ophthalmology, pathology, and dermatology.
الكلمات المفتاحية
deep learning
medical imaging
diagnostic accuracy
sensitivity and specificity
meta-analysis
external validation


