Skip to content

Federated Explainability for Network Anomaly Characterization: A Collaborative Effort by Ikerlan and Mondragon Unibertsitatea in relation the IDUNN Project

Within the IDUNN project, Mondragon Unibertsitatea and Ikerlan have developed a paper titled “Federated Explainability for Network Anomaly Characterization.” Machine learning (ML) based systems have demonstrated promising results for intrusion detection due to their capability to learn complex patterns. Particularly, unsupervised anomaly detection approaches offer practical advantages as they do not require labeling the training data, which is both costly and time-consuming. To address practical concerns further, there is a growing interest in adopting federated learning (FL) techniques as a modern ML model training paradigm for distributed settings (e.g., IoT), thereby tackling challenges such as data privacy, availability, and communication cost concerns.

However, the output generated by unsupervised models provides limited contextual information to security analysts at Security Operation Centers (SOCs). These models usually lack the means to explain why a sample was classified as anomalous or cannot differentiate between various types of anomalies, making it difficult to extract actionable information and correlate with other indicators. Additionally, ML explainability methods have received little attention in FL settings and present further challenges due to the distributed nature and data locality requirements.

This paper proposes a new methodology to characterize and explain the anomalies detected by unsupervised ML-based intrusion detection models in FL settings. It adapts and develops explainability, clustering, and cluster validation algorithms for FL settings to identify patterns in anomalous samples and recognize different threats across the entire network. The methodology is demonstrated on two network intrusion detection datasets containing real IoT malware, namely Gafgyt and Mirai, and various attack traces. The learned clustering results can be used to classify emerging anomalies, provide additional context that can be leveraged to gain more insights, and enable the correlation of anomalies with alerts triggered by other security solutions.

In Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses (RAID ’23). Association for Computing Machinery, New York, NY, USA, 346–365. https://doi.org/10.1145/3607199.3607234

 

Complete text in the link