Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

FlowSentry — Real-Time ML-Based IDS

FlowSentry is a real-time, Machine Learning-based Intrusion Detection System (IDS) that sniffs live network traffic, aggregates packets into sliding-window flow features, runs anomalies through a multi-model ML classifier gate, and invokes a LangGraph-driven response agent to suggest active firewall defense rules.


1. System Architecture

The following diagram illustrates the concurrent multi-staged data pipeline and agent reasoning loops:

flowchart TD
    subgraph NC["Network Capture (Userspace)"]
        A[Raw Network Traffic] -->|Packet Capture| B(Scapy AsyncSniffer)
        B -->|Queue Threadsafe Call| C[raw_queue]
    end

    subgraph FE["Feature Engineering"]
        C -->|Consumed by feature_loop| D(FlowWindow sliding state)
        D -->|5s time window features| E[feature_queue]
    end

    subgraph ML["ML Inference"]
        E -->|Consumed by inference_loop| F{IsolationForest Gate}
        F -->|Anomaly Score > -0.6| G[BENIGN - Ignored]
        F -->|Anomaly Score <= -0.6| H(XGBoost Multiclass Predictor)
        H -->|Class Probabilities| I(SHAP TreeExplainer)
        I -->|SHAP Values + Threat Dict| J[threat_queue]
    end

    subgraph RA["Response Agent (LangGraph)"]
        J -->|Consumed by persist_loop| K[(flowsentry.db SQLite)]
        J -->|Triggers process_threat| L[ThreatReader Node]
        L --> M[Planner Node]
        M -->|Conf > 0.88 & Severe| N[Responder Node: block_suggest]
        M -->|Conf > 0.70| O[Responder Node: alert]
        M -->|Else| P[Responder Node: log]
        N --> Q
        O --> Q
        P --> Q
        Q -->|Persist reasoning chain| R[(flowsentry_audit.db SQLite)]
    end

    subgraph UI["Visual UI Dashboard"]
        K --> S
        R --> S
        Q -->|WebSocket Broadcast| S
    end
Loading

2. Setup & Installation

Prerequisites

  • Python 3.10+
  • Node.js 18+
  • Root/Admin Privileges (required for binding raw network sockets via Scapy on macOS/Linux).

Backend Installation

  1. Navigate to the repository root directory and install dependencies:
    pip install -r requirements.txt
  2. Set up the database directories and models folder (these are automatically generated upon training/startup if they do not exist):
    mkdir -p models

Frontend Installation

  1. Navigate to the frontend directory:
    cd frontend
  2. Install node dependencies:
    npm install

3. Running the System

Step 1: Train the Machine Learning Models

Before running the engine, train the IsolationForest and XGBoost models:

python3 models/train.py

This downloads a subset of traffic features, trains an IsolationForest anomaly gate exclusively on benign data (representing a zero-day detection gate), trains an XGBoost classifier on attack classes, generates the SHAP TreeExplainer, and saves the trained pickles to models/.

Step 2: Start the Frontend Server

From the frontend folder:

npm run dev

The React UI dashboard will spin up at http://localhost:5173/.

Step 3: Start the Backend Engine (with Sudo)

From the root directory:

sudo python3 main.py

Important

Running with sudo is mandatory on macOS/Linux to allow Scapy to bind to the physical interface (e.g. en0 or eth0) and capture raw Ethernet packets. If run without sudo, the engine will safely fallback into Passive Mode (disabling packet sniffing while keeping APIs, databases, and WebSocket loops active for E2E tests and demo attack scripts).


4. Demo Attack & Integration Tests

FlowSentry includes automated scripts to generate mock traffic, verify detection accuracy, and test dashboard feeds.

Run End-to-End Integration Tests

You can verify the entire pipeline (capture, features, ML inference, LangGraph agent routing, SQLite persistence, and metrics) without raw sockets by injecting packets directly into the queue:

python3 tests/test_e2e.py

This runs the full E2E test suite simulating DDoS, Port Scan, SSH Brute Force, and Mixed slow-ramp attacks, producing a 7/7 success check.

Simulate Network Attacks

While both the backend (sudo python3 main.py) and frontend (npm run dev) are running, execute the simulation script to generate active traffic:

  1. DDoS Attack simulation (rapid SYN flood to port 80):
    python3 demo_attack.py --type ddos --count 300
  2. Port Scan simulation (sweeping destination ports sequentially):
    python3 demo_attack.py --type portscan --count 150
  3. SSH Brute-Force simulation (rapid connections targeting port 22):
    python3 demo_attack.py --type bruteforce --count 30
  4. Mixed Slow-Ramp Attack (interleaves benign traffic with a gradually escalating SYN attack to test IsolationForest detection gates):
    python3 demo_attack.py --type mixed

Observe the live dashboard at localhost:5173 to see metric rates spike, WebSocket threat feeds update instantly, and SHAP explainability bar charts render.


5. Scaling to Production

Transitioning FlowSentry from a single-host local prototype to a high-throughput, enterprise-scale Network IDS (NIDS) requires several structural upgrades:

[ Lightweight Capture Agent (eBPF/XDP) ] ──► [ Kafka Topic: flow-logs ] ──► [ Consumer Group: Processing Workers ] ──► [ Redis / TSDB Cache ]

A. Distributed Sensor Capture (Kafka Event Streaming)

  • Problem: Running a single Python sniffer thread limits capture to a single network adapter and interface. High-traffic environments (10 Gbps+) will crash Python's event loop and lose packet feeds.
  • Solution: Decouple packet capture from feature extraction. Deploy lightweight sniffer agents (written in Go/Rust) across multiple network switches or host mirrors. These agents extract minimal flow logs or packet headers locally and publish them to an Apache Kafka or Redpanda cluster.
  • Benefits: Under a consumer group configuration, workers can scale horizontally to consume packets, partitioned by flow hash hash(src_ip, dst_ip) to guarantee that packets belonging to the same flow are always processed by the same worker, ensuring correct sliding-window feature aggregation.

B. Kernel-Level Capture (eBPF / XDP)

  • Problem: Userspace libraries like Scapy or libpcap copy raw ethernet frames from kernel-space to userspace via system calls. This context-switching incurs extreme CPU overhead, dropping up to 80% of packets under high-throughput conditions.
  • Solution: Implement packet capture using eBPF (Extended Berkeley Packet Filter) or XDP (eXpress Data Path) programs loaded directly into the network driver level inside the Linux kernel.
  • Benefits: eBPF extracts packet flags, sizes, and headers in-kernel and aggregates them using high-performance kernel maps (ring buffers). The userspace Python/Go engine simply reads aggregated flow arrays from kernel maps periodically, eliminating packet copies, reducing context-switch overhead, and enabling wire-speed packet analysis (10G/40Gbps) with near-zero CPU footprint.

C. Distributed State Management (Redis / TSDB)

  • Problem: Storing sliding-window flow states in an in-memory python dictionary (deque(maxlen=300)) is volatile, non-distributed, and limits window aggregation to a single local thread, risking Out-Of-Memory (OOM) exceptions.
  • Solution: Replace python collections with a distributed cache like Redis or a high-performance Time-Series Database (TSDB) like TimescaleDB.
  • Implementation: Write packet records to Redis Sorted Sets (ZSET), where the score is the packet timestamp. Use ZREMRANGEBYSCORE to evict packets older than the 5-second window.
  • Benefits: Centralizing state in Redis allows multiple stateless engine workers to query flow states concurrently, enabling horizontal scalability, state replication, and high-availability (resilience to worker node failures).

D. Model Drift & Continuous Retraining

  • Problem: Network traffic distributions are highly dynamic. Software updates, new browser behavior, or novel attack payloads cause statistical covariate shift (data drift), leading to false positives or missed anomalies.
  • Solution: Deploy an automated model monitoring and retraining loop (e.g., using MLflow and Airflow/Kubeflow).
  • Implementation:
    1. Monitor features for drift using statistics like Population Stability Index (PSI) or Wasserstein Distance.
    2. Log feedback labels when security operators approve/dismiss LangGraph block decisions.
    3. When feature drift or accuracy degradation crosses a threshold, trigger an automated DAG to retrieve fresh datasets from SQLite/Postgres logs, retrain IsolationForest/XGBoost, deploy in a "shadow" configuration to measure performance, and swap the production model dynamically.

About

Real-time network intrusion detection system using a two-stage ML pipeline (IsolationForest + XGBoost) with SHAP explainability and an AI agent for automated mitigation suggestions.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages