Unlocking the Power of Unsupervised Learning: A Systematic Review of Anomalous Traffic Detection in Network Flows

Wednesday 09 April 2025


The eternal quest for a foolproof way to detect malicious activity on computer networks has led researchers to explore innovative approaches, including unsupervised machine learning algorithms applied to network flow data. A recent systematic review of these methods and their corresponding datasets has shed light on the state of the art in this field.


Flow-based analysis is an attractive alternative to traditional packet-based methods for detecting anomalies in network traffic. This approach focuses on aggregating packets into flows, which are collections of packets sent between two hosts over a specific time period. By analyzing these flows, researchers can identify unusual patterns that may indicate malicious activity.


The review, conducted by a team of researchers from the University of León, analyzed 63 scientific articles and selected 13 for in-depth evaluation. The study found that autoencoders are the most commonly used unsupervised learning algorithms for detecting anomalies in network flows. These models learn to compress data while preserving its essential features, making it easier to identify deviations from normal behavior.


Another key finding was the diversity of datasets used in these studies. A total of 16 datasets were identified, including some specifically designed for IoT and industrial control systems. The majority of these datasets are synthetic, but several contain real-world traffic captured from honeypots or other sources.


One of the most widely used datasets is CIC-IDS-2017, which has been employed in four separate studies. This dataset consists of network flow data collected over a period of several months and includes both normal and malicious traffic.


The review also highlighted some limitations of these methods. For instance, many of the datasets are small or contain biased samples, which can affect the accuracy of the algorithms. Additionally, the quality of the data itself is often uncertain, as it may be incomplete or contain errors.


Despite these challenges, the researchers remain optimistic about the potential of unsupervised machine learning for detecting anomalies in network traffic. By continuing to develop and refine these methods, they hope to create more effective tools for identifying and mitigating cyber threats.


In the future, we can expect to see further research into the use of generative adversarial networks (GANs) for anomaly detection. These models have shown promise in other applications, such as image recognition and natural language processing, and could potentially be applied to network traffic data.


Ultimately, the fight against cybercrime requires a multifaceted approach that incorporates multiple techniques and technologies.


Cite this article: “Unlocking the Power of Unsupervised Learning: A Systematic Review of Anomalous Traffic Detection in Network Flows”, The Science Archive, 2025.


Machine Learning, Network Flow Data, Anomaly Detection, Unsupervised Learning, Autoencoders, Iot, Industrial Control Systems, Honeypots, Cic-Ids-2017, Cyber Threats, Gans


Reference: Alberto Miguel-Diez, Adrián Campazas-Vega, Claudia Álvarez-Aparicio, Gonzalo Esteban-Costales, Ángel Manuel Guerrero-Higueras, “A systematic literature review of unsupervised learning algorithms for anomalous traffic detection based on flows” (2025).


Leave a Reply