PAIS26 Conference Presentation

Automated Privacy-PreservingFederated Learning

For IoT Intrusion Detection

Nasreddine Slimani
Dr. Mostefa Bendjima University of Bechar

Abstract & Key Contributions

Core Concept

Traditional centralized ML requires sharing raw IoT traffic data, violating privacy and compliance requirements. This work presents a fully automated federated learning pipeline where IoT devices train locally and share only model parameters.

The system requires zero expert configuration, reducing deployment from 8+ hours to under 1 hour.

Why It Matters

The platform preserves privacy, supports extreme non-IID distributions, and maintains near-centralized performance for Mirai botnet detection, making it practical for healthcare-adjacent and security-sensitive environments where data confidentiality matters.

99.91%
Accuracy
< 1 hr
Setup Time
99.98%
Comm. Savings
0.004%
False Positives

Problem & Motivation

📊

Severe Non-IID Data

IoT deployments create extreme heterogeneity. Some devices see mostly attack traffic while others observe benign traffic. Standard federated learning often degrades under these conditions.

⚠️

Single-Class Clients

Clients with 100% single-class labels can break traditional local learning. This is a realistic edge case for IoT devices and was not handled by prior FL-IDS approaches.

⏱️

Manual Configuration Burden

Typical FL workflows require expert configuration for partitioning, tuning, orchestration, and evaluation, often taking more than eight hours for one deployment.

Four-Stage Automated Pipeline

1

Data Partitioning

The N-BaIoT dataset is partitioned automatically with a Dirichlet distribution (α = 0.5) to simulate realistic non-IID behavior across clients.

2

Client Configuration

Clients are configured automatically with balanced, skewed, and single-class distributions to stress-test the learning system.

3

Training Orchestration

A SAGA-based logistic regression model is trained locally, then aggregated through FedAvg weighted by sample counts.

4

Convergence & Evaluation

The pipeline detects convergence automatically and stops early when 99.99% of final accuracy is reached, enabling single-round deployment.

Automated Client Configuration

Client Samples Distribution Status
Client 0 178,471 60% Benign / 40% Attack Balanced
Client 1 178,471 5% Benign / 95% Attack Skewed
Client 2 178,473 0% Benign / 100% Attack Single-Class
Total 535,415 Non-IID (α = 0.5) Automated

Experimental Results

Centralized vs. Federated Performance

99.97% Performance Retention

Key Security Metrics

Evaluated on 959,496 test samples

99.94%
Attack Detection
840,425 True Positives
0.004%
False Positive Rate
Only 5 False Alarms
100%
Benign Recall
Zero valid traffic blocked, 118,574 True Negatives

Confusion Matrix Summary

118,574
True Negatives
5
False Positives
492
False Negatives
840,425
True Positives

Communication Efficiency

Total Data Transfer Required

Centralized (Raw Data)
469.76
MB
Automated FL (Params)
108.75
KB
4,423× Smaller Payload

1-Round Convergence

The automated convergence detector stops training after Round 1 when the threshold is met, supporting fast deployment on constrained infrastructure.

💰

99.98% Savings

Only 108.75 KB are transferred for the federated process, compared with about 470 MB for centralized raw-data collection.

Analysis & Insights

🎯 Simple but Distinctive Patterns

Mirai botnet traffic exhibits strong statistical signatures. Because of this, a lightweight SAGA logistic regression model with only 232 parameters achieves strong performance without complex deep learning.

🌐 Cross-Client Knowledge Transfer

Most notable result: Client 2 contains only attack labels, yet still reaches 99.91% accuracy because aggregation transfers benign-class information from the other clients.

📦 Small Model, Big Practicality

The model size is only 1.81 KB, which makes the approach highly suitable for bandwidth-limited and embedded IoT settings.

🧠 Automation as the Main Contribution

The paper’s value is not only strong accuracy. It also removes repetitive manual FL setup steps, which is critical for real deployments and easier adoption by practitioners.

Conclusion & Future Directions

Five Proven Outcomes

  • 1. Automated Deployment: Setup reduced from 8+ hours to under 1 hour.
  • 2. Near-Centralized Accuracy: 99.97% performance retention with only a 0.03% trade-off.
  • 3. Extreme Heterogeneity Support: 100% single-class clients can still learn effectively.
  • 4. Minimal Communication: 4,423× lower payload with single-round deployment.
  • 5. Production-Ready Security: 99.94% detection and 0.004% false positives.

Future Research

  • Expand to multi-attack classification scenarios
  • Deploy on Raspberry Pi and ESP32 hardware
  • Integrate automated differential privacy parameter selection
  • Reduce the remaining 0.03% accuracy gap