Machine Learning for Network Intrusion Detection: An Analytical Review of Methods, Benchmark Datasets, and the Gap Between Reported and Operational Performance
DOI:
https://doi.org/10.54878/sn31gk23Keywords:
intrusion detection, machine learning, deep learning, base-rate fallacy, class imbalance, benchmark datasets, adversarial machine learningAbstract
Machine learning has become the dominant research approach to network intrusion detection, and the literature routinely reports detection accuracies exceeding 99% on standard benchmarks. This analytical review argues that such figures, taken at face value, are systematically misleading, and it uses the quantitative tools of detection theory to explain why. We define the confusion-matrix metrics, accuracy, precision, recall, the false-positive rate, the F1 score, and the Matthews correlation coefficient, and show that accuracy is nearly uninformative on the extreme class imbalance that characterizes real network traffic, where attacks may constitute a fraction of a percent of all flows. We then formalize the base-rate problem: because precision depends on the prior probability of attack, an intrusion detector with a 99% detection rate and a 1% false-positive rate achieves a precision below 10% when attacks are rarer than one in a hundred, so that most alarms are false. We survey the taxonomy of classical and deep-learning methods and the benchmark datasets, KDD99 and NSL-KDD, ISCX2012, UNSW-NB15, and CICIDS2017/CSE-CIC-IDS2018, on which they are evaluated, and we synthesize the reported performance, which is uniformly high. Against this we set the evidence that such performance does not transfer: benchmark datasets are dated and biased, models overfit their idiosyncrasies, concept drift erodes accuracy over time, and adversarial manipulation can defeat detectors that score near-perfectly offline. We conclude that the field's central problem is not raising benchmark accuracy, which is effectively saturated, but closing the gap between reported and operational performance through realistic evaluation, base-rate-aware metrics, robustness to drift and adversaries, and reproducible reporting. The review is methodological and is weighted toward network-based, flow-level detection.References
Abbadi, D. (2025). Cyber threats and risk mitigation strategies for cloud systems and the Internet of Things. International Journal of Information & Digital Security, 3(2).
Almalki, S. M., & Abdelmajeed, N. T. (2023). A new intelligent model for phishing web sites detection. International Journal of Information & Digital Security, 1(1).
Axelsson, S. (2000). The base-rate fallacy and the difficulty of intrusion detection. ACM Transactions on Information and System Security, 3(3), 186–205. https://doi.org/10.1145/357830.357849
Buczak, A. L., & Guven, E. (2016). A survey of data mining and machine learning methods for cyber security intrusion detection. IEEE Communications Surveys & Tutorials, 18(2), 1153–1176. https://doi.org/10.1109/COMST.2015.2494502
Ferrag, M. A., Maglaras, L., Moschoyiannis, S., & Janicke, H. (2020). Deep learning for cyber security intrusion detection: Approaches, datasets, and comparative study. Journal of Information Security and Applications, 50, 102419. https://doi.org/10.1016/j.jisa.2019.102419
Ghurab, M., Gaphari, G., Alshami, F., Alshamy, R., & Othman, S. (2021). A detailed analysis of benchmark datasets for network intrusion detection system. Asian Journal of Research in Computer Science, 7(4), 14–33. https://doi.org/10.9734/ajrcos/2021/v7i430185
Hakke, D., et al. (2025). Performance evaluation of machine learning-based intrusion detection using NSL-KDD, UNSW-NB15 and CICIDS2017 datasets. International Journal of Applied Mathematics.
He, K., Kim, D. D., & Asghar, M. R. (2023). Adversarial machine learning for network intrusion detection systems: A comprehensive survey. IEEE Communications Surveys & Tutorials, 25(1), 538–566. https://doi.org/10.1109/COMST.2022.3233793
Hozouri, A., et al. (2025). A comprehensive survey on intrusion detection systems with advances in machine learning, deep learning and emerging cybersecurity challenges. Discover Artificial Intelligence, 5, 1–35. https://doi.org/10.1007/s44163-025-00243-7
Liu, H., & Lang, B. (2019). Machine learning and deep learning methods for intrusion detection systems: A survey. Applied Sciences, 9(20), 4396. https://doi.org/10.3390/app9204396
Moustafa, N., & Slay, J. (2015). UNSW-NB15: A comprehensive data set for network intrusion detection systems (UNSW-NB15 network data set). In Proceedings of the 2015 Military Communications and Information Systems Conference (MilCIS) (pp. 1–6). IEEE. https://doi.org/10.1109/MilCIS.2015.7348942
Sharafaldin, I., Lashkari, A. H., & Ghorbani, A. A. (2018). Toward generating a new intrusion detection dataset and intrusion traffic characterization. In Proceedings of the 4th International Conference on Information Systems Security and Privacy (ICISSP) (pp. 108–116). https://doi.org/10.5220/0006639801080116
Shiravi, A., Shiravi, H., Tavallaee, M., & Ghorbani, A. A. (2012). Toward developing a systematic approach to generate benchmark datasets for intrusion detection. Computers & Security, 31(3), 357–374. https://doi.org/10.1016/j.cose.2011.12.012
Sommer, R., & Paxson, V. (2010). Outside the closed world: On using machine learning for network intrusion detection. In Proceedings of the 2010 IEEE Symposium on Security and Privacy (pp. 305–316). https://doi.org/10.1109/SP.2010.25
Tavallaee, M., Bagheri, E., Lu, W., & Ghorbani, A. A. (2009). A detailed analysis of the KDD CUP 99 data set. In Proceedings of the 2009 IEEE Symposium on Computational Intelligence for Security and Defense Applications (CISDA) (pp. 1–6). https://doi.org/10.1109/CISDA.2009.5356528
Vinayakumar, R., Alazab, M., Soman, K. P., Poornachandran, P., Al-Nemrat, A., & Venkatraman, S. (2019). Deep learning approach for intelligent intrusion detection system. IEEE Access, 7, 41525–41550. https://doi.org/10.1109/ACCESS.2019.2895334