Machine Learning for Network Intrusion Detection: An Analytical Review of Methods, Benchmark Datasets, and the Gap Between Reported and Operational Performance

Nathan Lee
International Journal of Information & Digital Security 30 Jul 2026 675 مشاهدة

المستخلص

Machine learning has become the dominant research approach to network intrusion detection, and the literature routinely reports detection accuracies exceeding 99% on standard benchmarks. This analytical review argues that such figures, taken at face value, are systematically misleading, and it uses the quantitative tools of detection theory to explain why. We define the confusion-matrix metrics, accuracy, precision, recall, the false-positive rate, the F1 score, and the Matthews correlation coefficient, and show that accuracy is nearly uninformative on the extreme class imbalance that characterizes real network traffic, where attacks may constitute a fraction of a percent of all flows. We then formalize the base-rate problem: because precision depends on the prior probability of attack, an intrusion detector with a 99% detection rate and a 1% false-positive rate achieves a precision below 10% when attacks are rarer than one in a hundred, so that most alarms are false. We survey the taxonomy of classical and deep-learning methods and the benchmark datasets, KDD99 and NSL-KDD, ISCX2012, UNSW-NB15, and CICIDS2017/CSE-CIC-IDS2018, on which they are evaluated, and we synthesize the reported performance, which is uniformly high. Against this we set the evidence that such performance does not transfer: benchmark datasets are dated and biased, models overfit their idiosyncrasies, concept drift erodes accuracy over time, and adversarial manipulation can defeat detectors that score near-perfectly offline. We conclude that the field's central problem is not raising benchmark accuracy, which is effectively saturated, but closing the gap between reported and operational performance through realistic evaluation, base-rate-aware metrics, robustness to drift and adversaries, and reproducible reporting. The review is methodological and is weighted toward network-based, flow-level detection.

الكلمات المفتاحية

intrusion detection machine learning deep learning base-rate fallacy class imbalance benchmark datasets adversarial machine learning

هل أنت مستعد للنهوض بالبحث معنا؟

انضم إلى مجتمع متنامٍ من العلماء والباحثين، أو أرسل ورقتك البحثية للنشر المحكّم مع باحثي الإمارات.