Evaluating Machine Learning-Based IoT Device Identification Models for Security Applications

Eman Maali

Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · IoT Security

Overview

In an increasingly interconnected world, the proliferation of Internet of Things (IoT) devices presents both convenience and significant security challenges. As these devices become ubiquitous in homes, enterprises, and industrial settings, the ability to accurately identify them on a network is paramount for effective security management. This talk, presented by Eman Maali, a collaborative effort between Imperial College London and Georgia Tech, addresses a critical gap in current IoT security practices: the real-world practicality of machine learning (ML)-based device identification models. The core problem highlighted is the scenario where a network operator needs to identify all IoT devices affected by a newly discovered vulnerability to apply security patches efficiently. The accuracy and robustness of the underlying identification solution are thus of utmost importance.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction and real-world IoT identification challenge
  2. 3:00 Defining a practical model with three key attributes
  3. 4:20 Methodology for evaluating ML-based IoT identification
  4. 6:00 Detailed experimental setup for practical model evaluation
  5. 7:45 Summary of key findings on model degradation

Evaluating Machine Learning-Based IoT Device Identification Models for Security Applications

Speakers: Eman Maali

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=y04_a0uTDIM

Overview

In an increasingly interconnected world, the proliferation of Internet of Things (IoT) devices presents both convenience and significant security challenges. As these devices become ubiquitous in homes, enterprises, and industrial settings, the ability to accurately identify them on a network is paramount for effective security management. This talk, presented by Eman Maali, a collaborative effort between Imperial College London and Georgia Tech, addresses a critical gap in current IoT security practices: the real-world practicality of machine learning (ML)-based device identification models. The core problem highlighted is the scenario where a network operator needs to identify all IoT devices affected by a newly discovered vulnerability to apply security patches efficiently. The accuracy and robustness of the underlying identification solution are thus of utmost importance.

The presentation delves into a rigorous evaluation of existing ML-based IoT device identification solutions, questioning their suitability for practical, real-world deployments. While 70% of current identification solutions rely on machine learning, they often struggle with adaptability and spatio-temporal validation across diverse environments and network conditions. Maali and her team define "practicality" through a set of attributes—mode of operation, transferability, and observability—and systematically test how well current models perform against these criteria. The findings reveal significant performance degradations, underscoring the urgent need for more robust and reliable identification mechanisms to safeguard against emerging IoT threats.

Background

▶ Watch: Introduction and real-world IoT identification challenge (0:00)

The landscape of IoT device identification solutions can broadly be categorized into static, rule-based, and machine learning approaches. Static and rule-based methods often face deployment challenges, requiring frequent regeneration and manual updates. In contrast, machine learning-based solutions, which constitute approximately 70% of current approaches, promise greater adaptability. However, as the research highlights, these ML models frequently encounter issues related to adaptability and spatial-temporal validation, struggling to maintain performance when faced with changes in environment or network conditions. This discrepancy between academic promise and real-world performance forms the crux of the problem addressed by Maali's work.

To systematically evaluate the practical utility of these models, the research first established a comprehensive definition of a practical model. Drawing from a literature review, a practical model is characterized by its ability to ensure robust and reliable IoT device identification across diverse operational modes, deployment environments, and network conditions. From this definition, three critical attributes were distilled for evaluation:

  1. Mode of Operation: This refers to the different states an IoT device can be in, such as idle (low activity, waiting for input), active (performing tasks, user interaction), setup mode, or reset mode. A practical model should generalize across these varied operational states, as device behavior and network traffic patterns can differ significantly between modes.
  2. Transferability: This attribute assesses a model's ability to maintain performance when deployed in different contexts from where it was trained.
  • Spatial Transferability examines if a model trained on a specific network, ISP, router, or geographical location (e.g., UK lab) can generalize effectively to another (e.g., US lab).
  • Temporal Transferability investigates a model's stability over time, evaluating if a model trained at one point remains robust when tested weeks or months later (e.g., after one week, one year).
  1. Observability: This attribute focuses on how well a model performs under varying network conditions. This includes scenarios like packet loss, packet delay, or traffic sampling (where only a subset of network traffic is observed by the model). A practical solution must be robust to these common real-world network imperfections.

Prior research often focused on generalization, cost metrics, and ethical considerations. However, for the specific context of IoT security and real-world deployment, the most common and relevant attributes identified were generalization, robustness, and stability over time, which directly map to the defined practicality attributes. The fundamental premise of this research is to rigorously test whether existing ML-based IoT identification solutions can genuinely deliver on these practical requirements for network operators.

Key Findings

▶ Watch: Defining a practical model with three key attributes (3:00)

The comprehensive evaluation, spanning 140 distinct experiments across the three defined practicality attributes, revealed significant and often alarming performance degradations in existing machine learning-based IoT device identification models. While some level of degradation was anticipated, the extent of the observed reduction in accuracy was notably higher than expected.

The key findings are summarized as follows:

  • Mode of Operation Degradation: Devices exhibit distinct behavioral shifts when operating in idle versus active modes. Models trained on a mix of these modes often struggle to accurately classify devices when they are predominantly in one state or the other. This behavioral divergence significantly reduces identification performance, highlighting that a "one-size-fits-all" training approach across modes is insufficient.
  • Spatial Transferability Degradation: The ability of models to generalize across different geographical locations and network environments showed substantial decline. Performance degradation in spatial transferability ranged from 7.5% up to a staggering 74%. This indicates that a model trained in one region (e.g., the UK) may be largely ineffective when deployed in another (e.g., the US), even with efforts to keep testbed configurations identical.
  • Temporal Transferability Degradation: The accuracy of identification models deteriorates significantly over time. Performance degradation typically begins after just one week and worsens considerably, reaching up to 86% after a year. This underscores the dynamic nature of IoT device behavior and network interactions, rendering static models quickly obsolete.
  • Observability Degradation (Sampling): When models were trained on full network traffic but tested against sampled traffic (e.g., 100,000 or 5,000 samples), the performance was severely impacted. On average, the use of sampled traffic reduced identification performance by 70.90%. This is a critical finding for real-world scenarios where network operators often rely on traffic sampling to manage monitoring overhead.

These findings collectively demonstrate that current ML-based IoT device identification solutions, as typically presented in research, are far from "plug-and-play" for real-world security applications. Their inherent fragility across various operational contexts, geographical boundaries, timeframes, and network conditions poses a significant challenge for network operators aiming to secure their IoT deployments.

Technical Deep Dive

▶ Watch: Methodology for evaluating ML-based IoT identification (4:20)

The methodology employed in this study was meticulously designed to ensure a rigorous and fair evaluation of existing ML-based IoT device identification models. This involved careful consideration of three main components: data sets, models, and features.

For data collection, the researchers created a unique dataset by establishing identical testbeds in two distinct geographical locations: a lab in the UK and another in the US. This dual-lab approach allowed for robust testing of spatial transferability. To ensure comparability, the team used identical IoT devices, aligned their modes of operation during data collection, and synchronized the collection times. This meticulous approach aimed to minimize confounding variables and enable rigorous evaluation and feature analysis.

Regarding model selection, the team undertook an extensive literature review, initially identifying over 200 papers related to IoT device identification. This pool was then filtered to focus specifically on ML-based solutions. From the most prominent and seminal works, ten papers were selected for in-depth evaluation. The reproduction of these models presented a significant challenge: two models had public, runnable code, one required minor modifications, but seven had to be implemented from scratch based on their original paper descriptions. Furthermore, four papers lacked sufficient information regarding their feature extraction methods, rendering them unsuitable for reproduction and thus excluded from the final set of ten evaluated models. This highlights a broader issue in academic research reproducibility.

For feature extraction, the study strictly adhered to the exact methods described in each selected paper. This was crucial to ensure that the reproduced baselines were as close as possible to the original research results, allowing for a fair comparison and analysis of the models' performance under the defined practical attributes. Each model's baseline performance was established, and then subjected to a series of experimental evaluations across the three practicality attributes:

  • Mode of Operation: Models were trained using a dataset comprising both idle and active device samples. They were then tested against three scenarios: devices exclusively in idle mode, exclusively in active mode, and a mix of both.
  • Transferability:
  • Spatial Transferability was evaluated by training a model in one lab (e.g., UK) and testing it against data collected from the other lab (e.g., US), and vice versa.
  • Temporal Transferability involved training a model at a specific point in time and then testing its performance against data collected at various later intervals, ranging from one week up to 52 weeks (one year).
  • Observability: To assess robustness against network conditions, models were trained on full, unsampled network traffic. Their performance was then tested against datasets where traffic was sampled at different rates, specifically 100,000 samples and 5,000 samples, simulating scenarios where network operators might sample traffic to reduce monitoring overhead.

In total, 140 experimental evaluations were conducted, providing a comprehensive assessment of the practical attributes.

To illustrate the nature of the performance degradation, the talk delved into a specific example related to the mode of operation, using a model from the "MED paper." This particular model utilized LightGBM as its classification algorithm and relied on features such as incoming packet count, incoming byte count, TCP flags, port numbers, and IP addresses. When this model was trained on a mix of active and idle samples and then tested, it exhibited random classification behavior for devices solely in idle or active modes, yet performed "fairly good" when classifying devices in a mixed-mode environment.

To understand this phenomenon, the researchers analyzed the underlying data. They found that IoT devices tend to remain in idle mode for significantly longer periods than active mode. Furthermore, in active mode, devices exhibit a much wider variety of traffic distributions. This behavioral asymmetry causes models trained on aggregated data to struggle when confronted with the distinct, singular patterns of idle or the highly varied patterns of active states. The recommendation for network operators and researchers is to consider training models specifically on idle mode data, collected at different times, for more reliable predictions, given that devices spend most of their time in this state.

To address the observed performance degradation, especially the limitations of existing features, the study shifted its focus from increasing model complexity (which researchers often prioritize) to improving feature engineering. They leveraged Explainable AI (XAI) techniques, specifically TSN (t-SNE) plots and SHAP (SHapley Additive exPlanations) values, to pinpoint the causes of degradation and identify more effective features. For the MED paper example, TSN plots initially showed no clear separation between device types, indicating poor feature representation. Further analysis using SHAP values revealed that a significant portion—33% of the features—had zero SHAP values, meaning they contributed nothing to the model's predictions.

This finding, corroborated by feature permutation analysis, led to a critical insight: researchers frequently rely on simple statistical features like mean and variance. These first-order features often fail to capture the subtle nuances and complexities within network traffic distributions. The recommendation is to avoid such simplistic features and instead incorporate higher-order features like entropy and kurtosis, which are better equipped to describe the shape and variability of data distributions, thereby capturing more insightful behavioral patterns.

The Q&A session further illuminated the challenges of spatial transferability. The degradation observed between UK and US labs, despite identical device setups, was attributed to several factors:

  1. Overlapping features: Some features were present in both environments but with subtly different distributions, confusing the models.
  2. Region-specific features: Certain destination ports or network behaviors were unique to one region (e.g., specific ports used only in the US), which could be helpful if learned correctly but also contribute to confusion if a model trained in one region encounters unseen patterns in another.
  3. Domain contact differences: The number and types of domains contacted by devices varied significantly between regions (e.g., more domains contacted in the US than the UK), leading to discrepancies when models were transferred.

For temporal transferability, the recommendation was clear: continuous model retraining and data updates are essential, ideally every one to two weeks. The choice of model can also influence long-term robustness; LightGBM models tend to capture longer-term patterns more effectively, while Random Forest models, while robust for shorter periods, may not generalize as well over extended durations as they are less prone to overfitting in the short term.

Demo / Proof of Concept

▶ Watch: Detailed experimental setup for practical model evaluation (6:00)

The talk focused on the rigorous evaluation methodology and the analytical results of existing machine learning-based IoT device identification models. It did not feature a live demonstration or a proof of concept of a novel identification system. Instead, the emphasis was on dissecting the performance and limitations of currently published approaches through extensive experimentation and data analysis.

Defensive Implications

▶ Watch: Summary of key findings on model degradation (7:45)

The findings of this research carry significant implications for network operators, security professionals, and IoT device manufacturers striving to secure their environments. The observed performance degradations highlight critical vulnerabilities in relying on generic or out-of-the-box ML-based IoT identification solutions.

  1. Continuous Model Retraining is Essential: The substantial degradation in temporal transferability (up to 86% after a year, starting after just one week) mandates that ML models for IoT device identification cannot be static. Network operators must implement strategies for continuous model retraining and data updates, ideally on a weekly or bi-weekly basis. This ensures that the models remain current with evolving device behaviors, firmware updates, and network traffic patterns.
  2. Beware of Spatial Limitations: The significant performance drop (7.5% to 74%) in spatial transferability means that a model trained in one geographical region or network environment might perform poorly in another. Organizations with distributed operations or those deploying solutions globally cannot simply transfer models. Regional-specific training datasets and models, or highly generalized feature engineering that accounts for regional variations (e.g., different destination ports, domain contacts), are crucial.
  3. Feature Engineering Over Model Complexity: The research strongly suggests that the focus for improving identification accuracy should shift from increasingly complex ML models to more sophisticated feature engineering. Defensores should prioritize solutions that leverage higher-order features like entropy and kurtosis instead of relying on simple mean and variance. These richer features capture the nuances of device behavior more effectively, leading to more robust identification.
  4. Understand Device Operational Modes: Models must account for the distinct behavioral patterns of devices in different modes of operation (idle, active, setup). Training models on a generic mix of modes without proper consideration can lead to random classification in specific states. Network operators should consider training and deploying models optimized for the predominant device state (e.g., idle mode) or multi-modal models that can adapt to context.
  5. Impact of Network Conditions: The 70.90% performance reduction due to traffic sampling is a critical concern. While sampling is common for network monitoring, it severely compromises the accuracy of ML-based identification. Defenders need to evaluate the trade-off between monitoring overhead and identification accuracy. If high accuracy is required for security, full traffic capture might be necessary, or models specifically trained on sampled data must be developed and validated.
  6. Rigorous In-Situ Validation: Before deploying any ML-based IoT identification solution, security teams must perform rigorous, in-situ validation within their specific operational environment. Relying on reported academic performance without verifying against local conditions, device types, and network characteristics is a high-risk strategy.
  7. Consider Model Choice for Longevity: For long-term robustness, models like LightGBM might be preferable as they tend to capture longer-period patterns. For shorter-term, highly accurate detection, Random Forest models can be effective, provided they are regularly updated to prevent overfitting to transient patterns.

In essence, network operators must adopt a more critical and informed approach to ML-based IoT device identification. Blind trust in published results or generic solutions is insufficient; active management, continuous adaptation, and a deep understanding of the underlying data and features are paramount for building resilient IoT security postures.

Key Takeaways

  • ML-based IoT identification models face significant "practicality" challenges in real-world deployments, often failing to generalize across diverse operational modes, environments, and network conditions.
  • Performance degrades substantially over time and space: Models show up to 86% degradation after a year (temporal) and up to 74% degradation across different geographical locations (spatial).
  • Simple features are insufficient: First-order statistical features like mean and variance fail to capture critical nuances in device behavior; advanced features such as entropy and kurtosis are recommended for improved accuracy.
  • Continuous retraining is crucial: To maintain accuracy, ML models for IoT identification require frequent retraining and data updates, ideally every one to two weeks, due to the dynamic nature of IoT device behavior.
  • Network conditions severely impact accuracy: Traffic sampling, a common practice in network monitoring, can reduce identification model performance by an average of 70.90%, highlighting a critical trade-off between monitoring overhead and security.
  • Context matters for device modes: Models must account for distinct device behaviors in different operational modes (e.g., idle vs. active); training strategies should consider these behavioral shifts for robust classification.

About the Speaker(s)

Eman Maali is a researcher whose work focuses on the critical area of machine learning-based IoT device identification for security applications. The research presented in this talk is a collaborative effort between Imperial College London and Georgia Tech, highlighting an interdisciplinary approach to tackling complex challenges in cybersecurity. Her work, as demonstrated, emphasizes rigorous evaluation and practical applicability of theoretical models in real-world scenarios.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Competent, methodologically honest evaluation of ML-based IoT device fingerprinting — 140 experiments, dual-lab testbed, reproducibility-aware model selection. The numbers are real and the degradation findings are damning enough to be useful, but this is ultimately a benchmarking paper dressed as a talk: it diagnoses a known problem more rigorously than most, without delivering a solution.

Heather Calloway (CISO) — WEAK

Methodologically rigorous academic work that quantifies something real — ML-based IoT identification models degrade badly in the wild — but stops well short of the operator and governance implications that would make it actionable for anyone responsible for a security program. The findings are credible; the translation is missing.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025