IoT Intrusion Detection
Catching attacks on IoT networks, and asking what happens when the model gets poisoned.

Random Forest and XGBoost classifiers trained on the Edge-IIoTset dataset to detect attacks on IoT networks, plus an analysis of data poisoning and a proposed multi-agent security architecture.
The question
Industrial IoT networks are full of cheap devices that nobody patches. Can a classifier trained on network traffic tell an attack from normal behaviour, and how easy is it to fool?
Baseline: suspiciously perfect
I trained Random Forest and XGBoost on Edge-IIoTset, a public dataset of traffic from a real IoT and IIoT testbed. On the clean test set both models scored 1.0 on accuracy, precision, recall and F1, and they held that score at every tree count I tried (50, 100, 200 and 300 estimators).
That’s useful for deployment, because a small model runs fine on a resource-constrained gateway. It’s also a warning sign. A dataset this separable probably reflects a controlled lab more than real traffic, and a perfect score says little about attacks the model has never seen.
Poisoning the training data
Then I played the insider. Assume someone with access to the training pipeline edits the labels before training. I tested two attacks at 5%, 10%, 15%, 20% and 30% of the training set:
- Random flipping swaps labels in both directions.
- Targeted flipping relabels attacks as benign, which is exactly what a malicious insider wants: real attacks that look like normal traffic.
The models reacted very differently. XGBoost barely moved, keeping attack recall around 0.997 even with 30% of labels corrupted. Random Forest degraded steadily: its attack recall fell to about 0.90 at 30% random flipping, and the targeted attack pushed it below that sooner. In a security tool, every point of recall lost is a real attack that gets through.
Where it landed
Nothing inside either model notices that its training data was tampered with; it just learns whatever the labels say. So I proposed an autonomous multi-agent architecture that moves the defence into the pipeline: a poisoning-detection agent checks each versioned data snapshot, training only runs on signed manifests, new models go out as canary deployments with automatic rollback, and a drift agent triggers retraining when live behaviour shifts.
MSc coursework for Cyber Security and AI Case Studies at Anglia Ruskin University.


