Skip to content

Lightweight Federated Intrusion Detection for Campus IoT Networks on CICIoT2023

  • 13 slides
  • 15 viva questions
  • 5 modules
  • Code included

@federated-iot-intrusion-detectionUpdated Oct 2026

Centralised vs FedAvg training of a small MLP/1D-CNN across simulated campus gateways, with non-IID splits and Raspberry Pi latency

M.Tech / M.E., Computer Science & Engineering · Sem 4 · Advanced · 24 weeks · Solo

More info
Level
Advanced · 24 weeks · Solo
Relevant for
Karnataka
Common at
Visvesvaraya Technological University, Anna University, JNTU Hyderabad
Syllabus
VTU M.Tech 2022 Scheme · 22SXX41 Project Work Phase-2 · Semester 4
Tech stack
  • Python
  • PyTorch
  • Flower (industry-standard FL framework)
  • scikit-learn
  • pandas
  • ONNX Runtime
  • Raspberry Pi 4
For educational purposes only

Unlock this project

Full PPT + speaker notes, source code and setup steps, READMEFIRST, instructions and all 15 viva answers.

One-time. No subscription, no auto-renew, no drama.

Project packs

Credits never expire and work on any project. Use one here, save the rest for your friend who “will pay you back”.

  1. Pinned

    1 min

    Overview

    This dissertation studies whether a lightweight intrusion detection model can be trained collaboratively across the IoT gateways of an engineering campus without moving raw traffic to a central server. A typical Karnataka campus has separate network segments for hostels, laboratories, the library, CCTV and building management, each with cameras, smart plugs, access-control readers and sensors. Pooling all their traffic for training raises privacy, bandwidth and administrative concerns; federated learning (FL) lets each gateway train locally and share only model updates.

    Using the public CICIoT2023 dataset from the Canadian Institute for Cybersecurity, I build a centralised baseline and a federated version (FedAvg) of two compact models — a multilayer perceptron and a 1D-CNN — and compare them under IID and non-IID partitions that mimic gateways seeing very different attack mixes. The evaluation reports macro-F1, per-class recall, communication rounds to reach a target F1, bytes transmitted, model size and inference latency on a Raspberry Pi 4, the kind of device that could sit at a gateway.

    The implementation uses Python and PyTorch, with the Flower framework as an industry-standard extra for FL simulation. The work extends the Computer Networks & IoT Lab (22SCNL17) and applies the Research Methodology & IPR course (22RMI16) for experimental design, threats to validity and paper writing. All results are produced by the student's own runs; this bundle gives the protocol and empty results tables, not pre-filled numbers.

    Syllabus alignment

    VTU · M.Tech 2022 Scheme

    22SXX41 · Project Work Phase-2 · Semester 4 · 18 credits · CIE 100 + SEE 100 (viva)

    Subjects this project applies
    • 22SCNL17 Computer Networks & IoT Lab
    • 22RMI16 Research Methodology & IPR
    • 22SXX34 Project Work Phase-1 (literature survey and problem formulation)
    • Machine learning with Python
    How it is evaluated

    50 : 25 : 25 (report : presentation : Q&A)

    Also fits: Anna University M.E. Regulation 2021, JNTUH M.Tech R22.

    1 min read · 15 viva questions

  2. 2 min

    Synopsis

    Abstract

    IoT devices on campus networks are frequent targets of botnet recruitment, DDoS, reconnaissance and spoofing attacks. Machine-learning intrusion detection systems (IDS) perform well when trained on pooled traffic, but pooling is often impractical across departments that manage their own networks. This work evaluates federated learning for a lightweight IDS: each simulated campus gateway trains a small neural network on its local share of CICIoT2023, and a coordinator aggregates weights with FedAvg. I compare federated and centralised training under IID and label-skewed non-IID partitions at 2-class, 8-class and 34-class granularity, and measure detection quality, communication cost and edge inference latency.

    Introduction

    Federated learning, introduced as FedAvg by McMahan et al., keeps data on clients and exchanges only model parameters. For IoT security this is attractive: gateways already see local traffic, and the model can be small enough to run on a single-board computer. The difficulty is statistical heterogeneity — a hostel gateway may see mostly Mirai-style scans while a lab gateway sees brute-force attempts — which slows convergence and can hurt minority classes.

    Literature Gap

    • Many IDS studies report high accuracy on centralised splits of older datasets, where accuracy hides poor minority-class recall.
    • Federated IDS papers often use IID splits or a single granularity, and seldom report communication cost or on-device latency together.
    • Gap: a reproducible comparison of centralised vs FedAvg (with FedProx as an extension) on a recent large IoT dataset, under controlled non-IID skew, reporting macro-F1, per-class recall, rounds, bytes and Raspberry Pi latency for models under a fixed size budget.

    Proposed Work

    • Preprocessing pipeline with leakage-safe splits and class-balanced evaluation sets.
    • Two compact models (MLP and 1D-CNN) under a parameter budget.
    • FL simulation with 10 gateways, Dirichlet label skew (α ∈ {100, 1.0, 0.3, 0.1}) and partial participation.
    • Edge benchmark of exported models on a Raspberry Pi 4.

    Feasibility

    The dataset is public; PyTorch and Flower are open-source; simulation runs on a single GPU or a CPU workstation; a Raspberry Pi 4 costs a few thousand rupees and is available in most college IoT labs.

  3. 1 min

    Problem statement

    Campus IoT networks are segmented across departments, and each segment's gateway sees a different mix of benign and malicious traffic. A centralised IDS requires copying traffic from every segment to one server, which conflicts with data-minimisation expectations under India's Digital Personal Data Protection Act, 2023, consumes backbone bandwidth, and crosses administrative boundaries. Federated learning avoids moving raw data, but its behaviour on realistic, heavily imbalanced and non-IID IoT attack data — particularly for rare attack classes — and its communication and edge-compute costs are not well characterised.

    The problem addressed is: How close can a federated, lightweight IDS trained across simulated campus gateways come to a centralised model on CICIoT2023, in macro-F1 and per-class recall, as label skew increases — and at what communication cost and on-device latency?

  4. 1 min

    Objectives & scope

    1. 01Build a leakage-safe preprocessing pipeline for CICIoT2023 with 2-, 8- and 34-class label sets and stratified, class-balanced test sets.
    2. 02Train centralised MLP and 1D-CNN baselines under a fixed parameter budget and report macro-F1 and per-class recall.
    3. 03Simulate federated training across 10 campus gateways with FedAvg (and FedProx as an extension) using Flower.
    4. 04Quantify the effect of Dirichlet label skew and client participation rate on convergence, macro-F1 and minority-class recall.
    5. 05Measure communication cost (rounds and megabytes) needed to reach a target macro-F1.
    6. 06Export the best models to ONNX and measure size and inference latency on a Raspberry Pi 4.
    7. 07Write the results as a conference-style paper with threats to validity.

    Scope

    In scope

    • Offline experiments on the published CICIoT2023 CSV feature files; no live capture of campus traffic.
    • Two compact neural models plus a random-forest reference; FedAvg and FedProx aggregation.
    • Simulated 10-client federation on one machine; optional two-Pi real deployment for latency.
    • Metrics: macro-F1, per-class recall, confusion matrices, rounds and bytes, model size, latency.

    Out of scope

    • Secure aggregation, differential privacy and defences against poisoned clients (discussed as future work).
    • Payload inspection or deep-packet inspection.
    • Production deployment on the actual campus network.
  5. 2 min

    Methodology

    Research design: quantitative, controlled experimental study with repeated runs (5 seeds per configuration), following the experimental-design and reporting practices taught in 22RMI16.

    StageWeeksActivitiesOutput
    Phase-1 recap1–2Refine research questions and hypotheses from the Phase-1 literature surveyRQ document
    Data engineering3–5Download, merge, deduplicate, scale, create label maps and splitsClean parquet files, data card
    Centralised baselines6–8MLP, 1D-CNN, random forest; tuning on validation setBaseline table
    FL simulation9–14Flower clients/strategy, IID and Dirichlet splits, FedAvg then FedProx, participation 50–100%Convergence curves
    Edge benchmark15–16ONNX export, optional int8 quantisation, latency on Raspberry Pi 4Latency table
    Analysis17–19Statistical tests, error analysis, ablationsResults chapter
    Writing20–24Thesis, paper draft, plagiarism check, viva preparationSubmission

    Research questions

    • RQ1: Under IID partitioning, how does FedAvg macro-F1 compare with the centralised model?
    • RQ2: How do macro-F1 and minority-class recall degrade as Dirichlet α decreases?
    • RQ3: How many rounds and megabytes does FL need to reach a target macro-F1?
    • RQ4: Do the models meet an edge budget (target: < 1 MB, < 5 ms per flow batch on a Raspberry Pi 4)?

    Evaluation protocol

    • Split by capture file where possible, not by random rows, to reduce temporal leakage; fit scalers on training data only.
    • Hold out one global balanced test set; every configuration is evaluated on it.
    • Report mean ± standard deviation over 5 seeds; compare configurations with a paired Wilcoxon signed-rank test on seed-wise macro-F1.
    • Targets (hypotheses, not results): FedAvg within 2 macro-F1 points of centralised at α = 100; the gap widening at α = 0.1; FedProx reducing that gap.
  6. 2 min

    Architecture & tech stack

    • Python
    • PyTorch
    • Flower (industry-standard FL framework)
    • scikit-learn
    • pandas
    • ONNX Runtime
    • Raspberry Pi 4

    The experimental system has four parts: data pipeline, centralised trainer, federated simulation (Flower server + N gateway clients), and edge benchmark.

    flowchart TD
      A["CICIoT2023 CSV files"] --> B["Preprocess: dedupe, scale, label maps (2/8/34 classes)"]
      B --> C["Global balanced test set"]
      B --> D["Training pool"]
      D --> E["Centralised trainer: MLP, 1D-CNN, RF"]
      D --> F["Partitioner: IID or Dirichlet alpha"]
      F --> G1["Gateway client 1: Hostel"]
      F --> G2["Gateway client 2: Library"]
      F --> G3["Gateway clients 3 to 10: Labs, CCTV, Admin"]
      G1 --> S["Flower server: FedAvg or FedProx"]
      G2 --> S
      G3 --> S
      S --> G1
      S --> G2
      S --> G3
      E --> H["Evaluator: macro-F1, per-class recall"]
      S --> H
      C --> H
      S --> X["ONNX export"]
      X --> P["Raspberry Pi 4 latency benchmark"]

    Models (parameter budget ≈ 50k)

    • MLP: input (published CSV features) → 128 → 64 → classes; ReLU, dropout 0.2, batch norm.
    • 1D-CNN: features treated as a 1-D sequence; two Conv1d layers (32 and 64 filters, kernel 3), global average pooling, dense output.
    • Reference: random forest (scikit-learn), centralised only.
    • Loss: weighted cross-entropy; optimiser: Adam for centralised, SGD with momentum on clients.

    Federated setup

    • 10 clients, 1–5 local epochs, batch 256, 50–100 rounds, fraction fit 0.5 or 1.0.
    • FedAvg weights client updates by local sample count; FedProx adds a proximal term μ‖w − w_global‖² to the local loss.
    • Communication per round = 2 × model bytes × participating clients (upload + download), logged by the strategy.

    Results tables (to be filled from your runs)

    SettingModelGranularityMacro-F1 (mean ± sd)Min per-class recallRounds to targetMB transferred
    CentralisedMLP8-class——
    FedAvg α = 100MLP8-class
    FedAvg α = 0.1MLP8-class
    FedProx α = 0.1MLP8-class
    ModelParamsONNX size (KB)Pi 4 latency per 1,000 flows (ms)int8 size (KB)int8 macro-F1 drop
    MLP
    1D-CNN
  7. 5 modules

    Modules

    • Data Pipeline and Partitioner

      Loads CICIoT2023 CSV files in chunks, removes duplicates and infinite values, builds 2-, 8- and 34-class label maps, fits a standard scaler on training data only, writes parquet shards, and creates IID or Dirichlet label-skew partitions for ten simulated campus gateways with a reproducible seed.

    • Centralised Baselines

      Trains the MLP, 1D-CNN and random-forest reference on pooled training data with weighted cross-entropy, early stopping on a validation split, and logs macro-F1, per-class recall and confusion matrices on the global balanced test set.

    • Federated Simulation (Flower)

      Implements a Flower NumPy client that wraps the PyTorch model and a custom strategy extending FedAvg to log bytes per round and evaluate the global model centrally each round; FedProx is added through a proximal loss term on clients. Supports partial participation and varying local epochs.

    • Edge Benchmark

      Exports the global model to ONNX, applies optional dynamic int8 quantisation, and runs a timing harness with ONNX Runtime on a Raspberry Pi 4 measuring median and 95th-percentile latency per batch, model size and peak memory.

    • Analysis and Reporting

      Aggregates results across seeds, runs Wilcoxon signed-rank tests, plots convergence curves and per-class recall bars, and generates LaTeX or Word tables for the thesis and paper.

  8. Locked

    Presentation

    13 slides with speaker notes. The outline below is free; the bullets, notes and the generated .pptx unlock with the project.

    1. Lightweight Federated Intrusion Detection for Campus IoT
    2. Motivation
    3. Problem and Research Questions
    4. Related Work and Gap
    5. Dataset: CICIoT2023
    6. Methodology
    7. System Architecture
    8. Experimental Setup
    9. Results: Detection Quality
    10. Results: Communication and Edge
    11. Discussion and Threats to Validity
    12. Contributions
    13. Conclusion and Future Work

    Bullets, speaker notes and the .pptx download unlock with the project.

    Presentation is locked: 13 slides, Speaker notes, .pptx download.

  9. 1 min

    Future scope

    • Secure aggregation and differential privacy to protect client updates, with a privacy–utility curve.
    • Robust aggregation (median, trimmed mean) against poisoned or compromised gateways.
    • Personalised FL so each gateway fine-tunes for its own traffic mix.
    • Real deployment on campus Raspberry Pi gateways with live flow features from a lightweight extractor.
    • Continual learning to add new attack classes without forgetting old ones.
  10. 10 sources

    References

    1. CICIoT2023 dataset — Canadian Institute for Cybersecurity, University of New Brunswick
    2. E. C. P. Neto et al., CICIoT2023: A Real-Time Dataset and Benchmark for Large-Scale Attacks in IoT Environment, Sensors, 2023
    3. H. B. McMahan et al., Communication-Efficient Learning of Deep Networks from Decentralized Data, AISTATS 2017
    4. T. Li et al., Federated Optimization in Heterogeneous Networks (FedProx), MLSys 2020
    5. T.-M. H. Hsu, H. Qi & M. Brown, Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification, 2019
    6. D. J. Beutel et al., Flower: A Friendly Federated Learning Research Framework, 2020
    7. P. Kairouz et al., Advances and Open Problems in Federated Learning, Foundations and Trends in Machine Learning, 2021
    8. Flower Framework Documentation
    9. PyTorch Documentation
    10. ONNX Runtime Documentation

    Cite this bundle

    OnlyProjects. (2026). Lightweight Federated Intrusion Detection for Campus IoT Networks on CICIoT2023: M.Tech / M.E. Computer Science & Engineering project bundle [Educational resource]. https://onlyprojects.online/projects/mtech-cse-federated-iot-intrusion-detection

Slides, diagrams & files

13 slides. Titles are free; bullets, speaker notes and the .pptx unlock with the project.

  1. SLIDE 1

    Lightweight Federated Intrusion Detection for Campus IoT

  2. SLIDE 2

    Motivation

  3. SLIDE 3

    Problem and Research Questions

  4. SLIDE 4

    Related Work and Gap

  5. SLIDE 5

    Dataset: CICIoT2023

  6. SLIDE 6

    Methodology

  7. SLIDE 7

    System Architecture

  8. SLIDE 8

    Experimental Setup

  9. SLIDE 9

    Results: Detection Quality

  10. SLIDE 10

    Results: Communication and Edge

  11. SLIDE 11

    Discussion and Threats to Validity

  12. SLIDE 12

    Contributions

  13. SLIDE 13

    Conclusion and Future Work

Architecture diagram

1
flowchart TD
  A["CICIoT2023 CSV files"] --> B["Preprocess: dedupe, scale, label maps (2/8/34 classes)"]
  B --> C["Global balanced test set"]
  B --> D["Training pool"]
  D --> E["Centralised trainer: MLP, 1D-CNN, RF"]
  D --> F["Partitioner: IID or Dirichlet alpha"]
  F --> G1["Gateway client 1: Hostel"]
  F --> G2["Gateway client 2: Library"]
  F --> G3["Gateway clients 3 to 10: Labs, CCTV, Admin"]
  G1 --> S["Flower server: FedAvg or FedProx"]
  G2 --> S
  G3 --> S
  S --> G1
  S --> G2
  S --> G3
  E --> H["Evaluator: macro-F1, per-class recall"]
  S --> H
  C --> H
  S --> X["ONNX export"]
  X --> P["Raspberry Pi 4 latency benchmark"]

Files

Viva questions & answers

3 of 15 questions free. Explain each answer in your own words before you move on.

  1. General

    What is the novelty of your dissertation?

    The novelty is a reproducible, combined evaluation of centralised versus federated lightweight IDS on CICIoT2023 under controlled label skew, reporting macro-F1, minority-class recall, communication cost and Raspberry Pi latency together. Most prior federated IDS work reports only accuracy on IID splits.

  2. General

    What did you complete in Phase-1 and what is new in Phase-2?

    Phase-1 (22SXX34) contained the literature survey, research questions, dataset study and a preliminary centralised baseline. Phase-2 (22SXX41) adds the full federated simulation, non-IID experiments, FedProx, edge benchmark, statistical tests, thesis and paper draft.

  3. Concept

    How does FedAvg work?

    In each round the server sends the global model to selected clients, each client trains for a few local epochs on its own data, and the server averages returned weights in proportion to client sample counts. Raw traffic never leaves the gateway; only parameters are exchanged.

+12 more questions

They and the answers unlock with the project. Try answering the ones above yourself first. Your examiner will.

For educational purposes only. Use this bundle to understand how the project works, then build and write your own. Submitting it verbatim is between you, your conscience and your external examiner.