Diagnostic Performance of Deep Learning in Medical Imaging: An Analytical Review and Quantitative Synthesis of the Evidence

Authors

  • Wei Zhang Author
  • Okafor Chinedu Author

DOI:

https://doi.org/10.54878/mcr68n77

Keywords:

deep learning, medical imaging, diagnostic accuracy, sensitivity and specificity, meta-analysis, external validation

Abstract

Deep learning has produced a large body of studies reporting expert-level accuracy in the diagnostic interpretation of medical images, yet the headline numbers are easily misread. This analytical review examines the quantitative evidence with the tools of diagnostic-test methodology, asking not only how accurate deep-learning models are but what their reported accuracy means and how far it can be trusted. We define and apply the core metrics, sensitivity, specificity, the area under the receiver operating characteristic curve (AUC), Youden's index, likelihood ratios, and the diagnostic odds ratio, and we assemble reported pooled estimates from the principal meta-analyses across imaging domains. The best available synthesis of externally validated studies places the pooled sensitivity of deep-learning models at 87.0% and specificity at 92.5%, statistically indistinguishable from the 86.4% and 90.5% of health-care professionals assessed on the same samples, while domain-specific meta-analyses report AUCs from roughly 0.86 for respiratory imaging to above 0.93 for retinal disease. Behind these figures, however, lies a consistent methodological problem that our analysis foregrounds: extreme between-study heterogeneity, a scarcity of external validation and prospective testing, small and non-representative human comparators, poor adherence to reporting standards, and susceptibility to dataset shift and confounding, all of which tend to inflate apparent performance. We argue that the central analytical lesson is a gap between reported diagnostic accuracy and demonstrated clinical validity, and we set out the metrics, study designs, and reporting practices, including AI-specific reporting guidelines and prospective external validation, needed to close it. The review is quantitative and methodological in emphasis and is weighted toward classification tasks in radiology, ophthalmology, pathology, and dermatology.

References

Aggarwal, R., Sounderajah, V., Martin, G., Ting, D. S. W., Karthikesalingam, A., King, D., … Darzi, A. (2021). Diagnostic accuracy of deep learning in medical imaging: A systematic review and meta-analysis. npj Digital Medicine, 4, 65. https://doi.org/10.1038/s41746-021-00438-z

Alshamsi, M., & Alhammadi, N. (2025). Next-generation AI in neurodevelopment: Multi-omics applications from diagnosis to care. International Journal of Rehabilitation & Disability Studies, 1(2), 13–24. https://doi.org/10.54878/yrp30g15

Esteva, A., Kuprel, B., Novoa, R. A., Ko, J., Swetter, S. M., Blau, H. M., & Thrun, S. (2017). Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542, 115–118. https://doi.org/10.1038/nature21056

Gulshan, V., Peng, L., Coram, M., Stumpe, M. C., Wu, D., Narayanaswamy, A., … Webster, D. R. (2016). Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA, 316(22), 2402–2410. https://doi.org/10.1001/jama.2016.17216

Kelly, C. J., Karthikesalingam, A., Suleyman, M., Corrado, G., & King, D. (2019). Key challenges for delivering clinical impact with artificial intelligence. BMC Medicine, 17(1), 195. https://doi.org/10.1186/s12916-019-1426-2

Liu, X., Faes, L., Kale, A. U., Wagner, S. K., Fu, D. J., Bruynseels, A., … Denniston, A. K. (2019). A comparison of deep learning performance against health-care professionals in detecting diseases from medical imaging: A systematic review and meta-analysis. The Lancet Digital Health, 1(6), e271–e297. https://doi.org/10.1016/S2589-7500(19)30123-2

Liu, X., Cruz Rivera, S., Moher, D., Calvert, M. J., & Denniston, A. K. (2020). Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: The CONSORT-AI extension. Nature Medicine, 26(9), 1364–1374. https://doi.org/10.1038/s41591-020-1034-x

McKinney, S. M., Sieniek, M., Godbole, V., Godwin, J., Antropova, N., Ashrafian, H., … Shetty, S. (2020). International evaluation of an AI system for breast cancer screening. Nature, 577, 89–94. https://doi.org/10.1038/s41586-019-1799-6

Nagendran, M., Chen, Y., Lovejoy, C. A., Gordon, A. C., Komorowski, M., Harvey, H., … Maruthappu, M. (2020). Artificial intelligence versus clinicians: Systematic review of design, reporting standards, and claims of deep learning studies. BMJ, 368, m689. https://doi.org/10.1136/bmj.m689

Rajkomar, A., Dean, J., & Kohane, I. (2019). Machine learning in medicine. New England Journal of Medicine, 380(14), 1347–1358. https://doi.org/10.1056/NEJMra1814259

Sulthan, N., et al. (2025). The role of AI in early diagnosis, management, and clinical research of primary immunodeficiency diseases: Insights from GCC immunologists. International Journal of Applied Technology in Medical Sciences, 4(1).

Topol, E. J. (2019). High-performance medicine: The convergence of human and artificial intelligence. Nature Medicine, 25(1), 44–56. https://doi.org/10.1038/s41591-018-0300-7

van der Laak, J., Litjens, G., & Ciompi, F. (2021). Deep learning in histopathology: The path to the clinic. Nature Medicine, 27(5), 775–784. https://doi.org/10.1038/s41591-021-01343-4

Xue, P., Wang, J., Qin, D., Yan, H., Qu, Y., Seery, S., … Qiao, Y. (2022). Deep learning in image-based breast and cervical cancer detection: A systematic review and meta-analysis. npj Digital Medicine, 5, 19. https://doi.org/10.1038/s41746-022-00559-z

Yu, K.-H., Beam, A. L., & Kohane, I. S. (2018). Artificial intelligence in healthcare. Nature Biomedical Engineering, 2(10), 719–731. https://doi.org/10.1038/s41551-018-0305-z

Zech, J. R., Badgeley, M. A., Liu, M., Costa, A. B., Titano, J. J., & Oermann, E. K. (2018). Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study. PLoS Medicine, 15(11), e1002683. https://doi.org/10.1371/journal.pmed.1002683

Downloads

Published

2026-06-30

How to Cite

Zhang, W., & Chinedu, O. (2026). Diagnostic Performance of Deep Learning in Medical Imaging: An Analytical Review and Quantitative Synthesis of the Evidence. International Journal of Applied Technology in Medical Sciences, 5(1), 17-25. https://doi.org/10.54878/mcr68n77