Skip to content

Latest commit

Β 

History

17 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸ›‘οΈ ShadowKYC

Real-Time KYC Video Integrity & Deepfake Risk Detection System

ShadowKYC is a multi-layer video integrity validation system designed to identify indicators of deepfakes, replay attacks, presentation attacks, face manipulation, temporal inconsistencies, and environmental anomalies during Video KYC.

Instead of replacing existing KYC providers, ShadowKYC acts as an additional security and risk-analysis layer that can run alongside existing identity-verification systems.


🚨 Problem

Modern Video KYC systems can be exposed to attacks such as:

  • AI-generated or face-swapped videos
  • Replay attacks using previously recorded KYC footage
  • Frozen-frame or screen-replay attacks
  • Synthetic facial texture and frequency artifacts
  • Abnormal facial geometry
  • Identity drift during a session
  • Lip/mouth movement inconsistencies
  • Background and lighting manipulation
  • Overlay and screen-capture artifacts

Traditional identity verification may establish who a person claims to be, but a separate integrity layer is useful for determining whether the presented video itself appears trustworthy.


πŸ’‘ Solution

ShadowKYC analyzes a KYC video at multiple levels and combines the resulting signals into a single explainable session-level risk score.

                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚      KYC Video Stream    β”‚
                    β”‚   Live Camera / Upload   β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β”‚
                                 β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚   Frame Sampling & Face  β”‚
                    β”‚        Detection         β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β”‚
          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
          β”‚                      β”‚                      β”‚
          β–Ό                      β–Ό                      β–Ό
   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
   β”‚ Layer 1     β”‚        β”‚ Layer 2     β”‚        β”‚ Layer 3     β”‚
   β”‚ Temporal    β”‚        β”‚ Face Match  β”‚        β”‚ Texture /   β”‚
   β”‚ Liveness    β”‚        β”‚ & Identity  β”‚        β”‚ Frequency   β”‚
   β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜        β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜        β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
          β”‚                      β”‚                      β”‚
          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
                         β”‚                      β”‚
                         β–Ό                      β–Ό
                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                  β”‚ Layer 4     β”‚        β”‚ Layer 5     β”‚
                  β”‚ Geometry /  β”‚        β”‚ Lip-Sync    β”‚
                  β”‚ Pseudo-Depthβ”‚        β”‚ Integrity   β”‚
                  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜        β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
                         β”‚                      β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                    β”‚
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚                     β”‚
                         β–Ό                     β–Ό
                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                  β”‚ Layer 6     β”‚       β”‚ Layer 7     β”‚
                  β”‚ Environment │──────▢│ Meta Fusion β”‚
                  β”‚ Integrity   β”‚       β”‚ & Decision  β”‚
                  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜       β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
                                               β”‚
                                               β–Ό
                                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                                  β”‚ Session Risk Assessmentβ”‚
                                  β”‚ LOW / MEDIUM / HIGH    β”‚
                                  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

🧠 Seven-Layer Detection Architecture

ShadowKYC currently uses seven complementary analysis layers.

Layer 1 β€” Temporal Liveness

Module: app/modules/layer_1_temporal_liveness.py

Analyzes temporal facial behavior to identify signs of replay or unnatural movement.

Signals

  • MediaPipe FaceMesh landmarks
  • Eye Aspect Ratio (EAR)
  • Blink variation
  • Head pose / yaw / pitch
  • Temporal head movement
  • Frozen-frame detection
  • Perceptual hashing
  • Face mesh availability

Example indicators

FROZEN_FRAME_REPLAY
L1_UNNATURAL_HEAD_STABILITY
L1_NO_BLINK_VARIATION
L1_EXTREME_HEAD_POSE
L1_NO_FACE_MESH

Layer 2 β€” Face Match & Identity Continuity

Module: app/modules/layer_2_face_match.py

Maintains a rolling facial baseline and checks whether the identity representation remains consistent throughout the session.

Signals

  • HOG-based CPU-friendly face representation
  • Cosine distance
  • Rolling face baseline
  • Embedding drift
  • Sudden facial appearance changes

Example indicators

L2_FACE_SWAP_DETECTED
L2_IDENTITY_DRIFT
L2_SUDDEN_APPEARANCE_CHANGE

This layer is intended to identify identity continuity problems, rather than acting as a standalone identity-verification provider.


Layer 3 β€” Texture & Frequency Artifact Detection

Module: app/modules/layer_3_texture_artifact.py

Analyzes facial texture and frequency characteristics that may become abnormal after synthetic generation, manipulation, compression, or presentation attacks.

Signals

  • Laplacian variance
  • FFT high-frequency energy ratio
  • Simplified LBP variance
  • Face ROI texture analysis
  • Excessive smoothing
  • Frequency-domain anomalies

Current thresholds

BLUR_THRESHOLD      = 80.0
FFT_LOW_THRESHOLD   = 0.08
FFT_HIGH_THRESHOLD  = 0.45
LBP_LOW_THRESHOLD   = 20.0

Example indicators

L3_OVER_SMOOTH_FACE
L3_LOW_HF_ENERGY_AI_FACE
L3_GAN_GRID_ARTIFACT
L3_UNIFORM_SKIN_TEXTURE

The layer also records metrics such as:

l3_laplacian_var
l3_fft_hf_ratio
l3_lbp_var

Layer 4 β€” Geometry & Pseudo-Depth Integrity

Module: app/modules/layer_4_geometry.py

Uses facial landmarks and geometric relationships to identify abnormal facial proportions or temporal geometry changes.

Signals

  • MediaPipe FaceMesh
  • Eye-width ratio
  • Eye-to-nose ratio
  • Nose-to-chin ratio
  • Facial symmetry
  • Ratio drift across frames
  • Eye-width distortion

Current thresholds

SYMMETRY_THRESHOLD = 0.25
RATIO_DRIFT_THRESH = 0.15
HISTORY_LEN        = 20

Example indicators

L4_FACE_ASYMMETRY
L4_LANDMARK_RATIO_DRIFT
L4_EYE_WIDTH_DISTORTION
L4_NO_LANDMARKS

Layer 5 β€” Lip-Sync Integrity

Module: app/modules/layer_5_lipsync.py

Analyzes mouth movement over time to identify frozen or highly abnormal facial motion.

Signals

  • MediaPipe FaceMesh
  • Mouth Aspect Ratio (MAR)
  • Mouth movement variance
  • Optional audio RMS
  • Optional audio/MAR correlation

Current thresholds

MAR_FROZEN_THRESHOLD  = 0.0005
MAR_ERRATIC_THRESHOLD = 0.08
HISTORY_LEN           = 30
MIN_FRAMES_FOR_EVAL   = 10

Example indicators

L5_MOUTH_FROZEN_NO_SPEECH
L5_ERRATIC_MOUTH_MOVEMENT
L5_AUDIO_LIP_MISMATCH

Note: The current implementation primarily uses visual mouth-motion analysis. Audio correlation is optional and should not be described as a complete audio deepfake detector.


Layer 6 β€” Environment & Background Integrity

Module: app/modules/layer_6_environment.py

Analyzes the environment surrounding the face to detect inconsistencies that can occur during replay, overlays, screen capture, or manipulated video.

Signals

  • Color histogram drift
  • Bhattacharyya distance
  • Background motion
  • Frame-difference analysis
  • Edge shimmer / overlay halo
  • Background flickering

Current thresholds

HIST_DIFF_THRESHOLD     = 0.35
BG_MOTION_THRESHOLD     = 15.0
EDGE_SHIMMER_THRESHOLD  = 0.12
HISTORY_LEN             = 20

Example indicators

L6_SUDDEN_LIGHTING_SHIFT
L6_BACKGROUND_MOTION_DETECTED
L6_EDGE_SHIMMER_OVERLAY
L6_BACKGROUND_FLICKERING

Layer 7 β€” Meta Fusion & Explainable Decision

Module: app/modules/layer_7_meta_fusion.py

The final layer converts the outputs from the six detection layers into a unified risk score.

Each individual module produces a goodness score, where:

1.0 = clean / authentic signal
0.0 = suspicious signal

Risk contribution is calculated as:

risk contribution = 1 - layer score

Layer Weights

Layer Signal Weight
L1 Temporal Liveness 0.25
L2 Face Match 0.20
L3 Texture / Frequency 0.15
L4 Geometry 0.15
L5 Lip-Sync 0.10
L6 Environment 0.15
Total 1.00

Contradiction Rules

ShadowKYC also applies additional penalties when multiple suspicious signals contradict each other in meaningful ways.

Combination Additional Penalty
Unnatural head stability + background motion +0.10
Face swap + frozen-frame replay +0.15
Over-smooth face + eye-width distortion +0.10
Frozen mouth + background motion +0.12

This allows the final decision to consider cross-layer evidence, rather than treating every layer independently.


🎯 Risk Classification

The final risk score is classified into three levels.

Risk Score Classification Recommendation
< 0.35 🟒 LOW RISK APPROVE β€” Session appears authentic
0.35 – < 0.65 🟑 MEDIUM RISK REVIEW β€” Manual verification recommended
>= 0.65 πŸ”΄ HIGH RISK REJECT β€” High probability of fraud

Example output:

{
  "risk_score": 0.72,
  "risk_pct": 72,
  "classification": "HIGH_RISK",
  "recommendation": "REJECT – High probability of fraud. Do not proceed.",
  "top_reasons": [
    "Face swap indicators detected",
    "Frozen frame replay detected",
    "Cross-layer contradiction detected"
  ]
}

πŸ—οΈ System Architecture

                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚      Frontend        β”‚
                         β”‚ React + Vite         β”‚
                         β”‚ Tailwind CSS          β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                    β”‚
                         HTTP / WebSocket
                                    β”‚
                                    β–Ό
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚      FastAPI         β”‚
                         β”‚      Backend         β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                    β”‚
                                    β–Ό
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚     Orchestrator     β”‚
                         β”‚ app/core/orchestratorβ”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                    β”‚
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β”‚                      β”‚                      β”‚
             β–Ό                      β–Ό                      β–Ό
       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
       β”‚   L1     β”‚          β”‚   L2     β”‚          β”‚   L3     β”‚
       β”‚ Liveness β”‚          β”‚Face Matchβ”‚          β”‚ Texture  β”‚
       β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜          β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜          β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜
            β”‚                     β”‚                     β”‚
            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                  β”‚
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚                 β”‚
                         β–Ό                 β–Ό
                   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”      β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                   β”‚   L4     β”‚      β”‚   L5     β”‚
                   β”‚ Geometry β”‚      β”‚ Lip-Sync β”‚
                   β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜      β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜
                        β”‚                 β”‚
                        β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β”‚
                                 β–Ό
                           β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                           β”‚   L6     β”‚
                           β”‚Environmentβ”‚
                           β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜
                                β”‚
                                β–Ό
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚     L7      β”‚
                         β”‚ Meta Fusion β”‚
                         β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
                                β”‚
                                β–Ό
                       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                       β”‚ Risk + Evidence  β”‚
                       β”‚ + Recommendation β”‚
                       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ“ Project Structure

ShadowKYC2/
β”‚
β”œβ”€β”€ app/
β”‚   β”œβ”€β”€ core/
β”‚   β”‚   └── orchestrator.py
β”‚   β”‚
β”‚   β”œβ”€β”€ modules/
β”‚   β”‚   β”œβ”€β”€ layer_1_temporal_liveness.py
β”‚   β”‚   β”œβ”€β”€ layer_2_face_match.py
β”‚   β”‚   β”œβ”€β”€ layer_3_texture_artifact.py
β”‚   β”‚   β”œβ”€β”€ layer_4_geometry.py
β”‚   β”‚   β”œβ”€β”€ layer_5_lipsync.py
β”‚   β”‚   β”œβ”€β”€ layer_6_environment.py
β”‚   β”‚   └── layer_7_meta_fusion.py
β”‚   β”‚
β”‚   └── main.py
β”‚
β”œβ”€β”€ frontend/
β”‚   └── src/
β”‚       └── pages/
β”‚           └── DevDashboard.jsx
β”‚
β”œβ”€β”€ evidence/
β”œβ”€β”€ reports/
β”œβ”€β”€ temp_uploads/
β”‚
β”œβ”€β”€ tests/
β”‚
β”œβ”€β”€ requirements.txt
β”œβ”€β”€ implementation_plan.md
β”œβ”€β”€ supabase_schema.sql
β”œβ”€β”€ db.json
β”œβ”€β”€ run_server.bat
└── README.md

πŸ–₯️ User Interface

ShadowKYC supports two primary analysis workflows.

1. Live Session

The live workflow is designed for real-time KYC sessions.

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    SHADOW KYC β€” LIVE                       β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                              β”‚                             β”‚
β”‚       CAMERA FEED            β”‚       RISK ANALYSIS        β”‚
β”‚                              β”‚                             β”‚
β”‚     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”       β”‚    Risk Score: 42%         β”‚
β”‚     β”‚                β”‚       β”‚    MEDIUM RISK              β”‚
β”‚     β”‚     CAMERA     β”‚       β”‚                             β”‚
β”‚     β”‚      FEED      β”‚       β”‚    L1  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘           β”‚
β”‚     β”‚                β”‚       β”‚    L2  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘           β”‚
β”‚     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜       β”‚    L3  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘           β”‚
β”‚                              β”‚    L4  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘           β”‚
β”‚                              β”‚    L5  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘           β”‚
β”‚                              β”‚    L6  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘           β”‚
β”‚                              β”‚                             β”‚
β”‚ [ START ] [ STOP ]           β”‚    ⚠ Evidence             β”‚
β”‚                              β”‚    ⚠ Identity drift        β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                    FLAGGED FRAMES / EVENTS                 β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

The live interface can display:

  • Camera feed
  • Rolling risk score
  • Layer status
  • Detection flags
  • Flagged frames
  • Session timeline
  • Start / stop controls
  • Session report

🎞️ Post-Session Video Analysis

Recorded KYC videos can also be uploaded and analyzed after the session.

Upload Video
     β”‚
     β–Ό
Frame Extraction
     β”‚
     β–Ό
Face Detection
     β”‚
     β”œβ”€β”€ L1 Temporal Liveness
     β”œβ”€β”€ L2 Face Match
     β”œβ”€β”€ L3 Texture
     β”œβ”€β”€ L4 Geometry
     β”œβ”€β”€ L5 Lip-Sync
     └── L6 Environment
             β”‚
             β–Ό
        L7 Meta Fusion
             β”‚
             β–Ό
       Final Risk Report

The report can expose:

  • Overall risk score
  • Risk classification
  • Recommendation
  • Layer scores
  • Detection flags
  • Top reasons
  • Contradiction penalties
  • Evidence frames
  • Session-level metadata

πŸ”Œ API

The backend is implemented with FastAPI.

Important API workflows include:

/analyze-session
/upload-recording

The frontend communicates with the backend using HTTP requests and WebSocket-based live-session communication.


βš™οΈ Installation

Requirements

Recommended environment:

  • Python 3.10
  • Node.js 18+
  • npm
  • Git
  • Webcam for live testing
  • Linux / WSL / Windows supported depending on dependency configuration

1. Clone the Repository

git clone https://github.com/Nirmal0804/ShadowKYC2.git
cd ShadowKYC2

2. Create Python Virtual Environment

Windows

python -m venv .venv
.venv\Scripts\activate

Linux / macOS

python3 -m venv .venv
source .venv/bin/activate

3. Install Backend Dependencies

pip install -r requirements.txt

▢️ Running the Backend

uvicorn app.main:app --host 0.0.0.0 --port 8000 --reload

The backend will be available at:

http://localhost:8000

FastAPI documentation:

http://localhost:8000/docs

▢️ Running the Frontend

Open another terminal:

cd frontend
npm install
npm run dev

The frontend will normally be available through the Vite development server.


πŸ§ͺ Testing

Run the test suite from the project root:

python -m unittest discover -s tests -v

The project should be tested after changes to:

  • Individual detection layers
  • Orchestrator logic
  • API routes
  • Risk fusion
  • Frontend/backend communication

πŸ”„ Live Processing Flow

Browser Camera
      β”‚
      β–Ό
WebSocket
      β”‚
      β–Ό
FastAPI
      β”‚
      β–Ό
Orchestrator.start_session()
      β”‚
      β–Ό
process_single_frame()
      β”‚
      β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
      β”‚             β”‚
      β–Ό             β–Ό
 Face Detection   Frame Metadata
      β”‚
      β–Ό
 Seven-Layer Analysis
      β”‚
      β–Ό
 Rolling Session State
      β”‚
      β–Ό
 Live Risk Updates
      β”‚
      β–Ό
 Browser Dashboard
      β”‚
      β–Ό
finalize_live_session()
      β”‚
      β–Ό
 Session Report

πŸ“Š Explainability

ShadowKYC is designed around evidence-based risk scoring rather than a single black-box prediction.

Instead of returning only:

FAKE

the system can provide:

Risk Score: 72%

Classification:
HIGH_RISK

Evidence:
- Frozen frame replay indicator
- Identity drift
- Facial texture anomaly
- Cross-layer contradiction

Recommendation:
Manual investigation / reject according to KYC policy

This makes the system more useful for fraud analysts and compliance workflows.


πŸ›‘οΈ Security Model

ShadowKYC should be deployed as a security layer around a KYC workflow.

                 Customer
                    β”‚
                    β–Ό
              Video KYC Flow
                    β”‚
                    β–Ό
          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
          β”‚ Existing KYC      β”‚
          β”‚ Provider          β”‚
          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                    β”‚
                    β”‚
          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”
          β”‚    ShadowKYC      β”‚
          β”‚ Integrity Layer   β”‚
          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                    β”‚
                    β–Ό
          Risk + Evidence
                    β”‚
          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
          β–Ό                   β–Ό
       Approve              Review /
                           Reject

ShadowKYC does not need to replace the customer's existing identity-verification provider.


☁️ Deployment Considerations

For production deployment, separate application logic from persistent storage.

Recommended architecture:

                   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                   β”‚   Frontend      β”‚
                   β”‚ React / Vite    β”‚
                   β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                            β”‚
                            β–Ό
                   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                   β”‚ FastAPI Server  β”‚
                   β”‚ ShadowKYC       β”‚
                   β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           β”‚
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β”‚             β”‚             β”‚
             β–Ό             β–Ό             β–Ό
        MongoDB Atlas    S3/Object    Processing
        Metadata         Storage      Worker
             β”‚             β”‚
             β”‚             β”‚
             β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
                    β–Ό
              Session Reports

Large video and image files should preferably be stored in object storage rather than directly inside a database.

Database records should primarily contain:

  • Session ID
  • User/tenant reference
  • Analysis status
  • Risk score
  • Classification
  • Layer results
  • Evidence references
  • Timestamps
  • Storage object references

🧰 Technology Stack

Backend

  • Python
  • FastAPI
  • OpenCV
  • MediaPipe
  • NumPy
  • SciPy
  • PyZBar
  • WebSockets

Computer Vision

  • MediaPipe FaceMesh
  • OpenCV
  • HOG-based facial representation
  • FFT analysis
  • Laplacian variance
  • LBP-based texture analysis
  • Facial geometry analysis
  • Temporal signal analysis

Frontend

  • React
  • Vite
  • Tailwind CSS
  • Chart.js
  • Lucide icons

Planned / Deployment Infrastructure

  • MongoDB Atlas
  • Amazon S3
  • Cloud deployment
  • Private object storage
  • Presigned media URLs

MongoDB Atlas and S3 should only be considered production architecture components once they are actually integrated into the application.


πŸ“ˆ Performance Considerations

Real-time video analysis can be computationally expensive.

Potential optimizations include:

  • Process only selected frames per second
  • Resize face crops before analysis
  • Use conditional execution
  • Batch independent operations
  • Cache repeated calculations
  • Avoid unnecessary full-resolution processing
  • Use asynchronous processing where appropriate
  • Use ONNX / optimized inference where applicable
  • Use CPU-friendly algorithms for lightweight layers

A production system should balance:

Detection Accuracy
        ↕
Processing Latency
        ↕
Infrastructure Cost

πŸ§ͺ Evaluation Strategy

A robust evaluation should test ShadowKYC against multiple attack categories.

Authentic Sessions

  • Normal webcam sessions
  • Different lighting
  • Different backgrounds
  • Different camera qualities
  • Different face angles
  • Natural speaking and blinking

Attack Sessions

  • Replay videos
  • Frozen-frame attacks
  • Screen recordings
  • Face-swap videos
  • Synthetic/deepfake videos
  • Overlay attacks
  • Manipulated lighting/background
  • Low-quality compressed videos

Recommended evaluation metrics:

Precision
Recall
F1 Score
False Positive Rate
False Negative Rate
ROC-AUC
Detection Latency
Per-layer contribution
Session-level accuracy

⚠️ Limitations

ShadowKYC should be treated as a risk-analysis system, not an infallible deepfake detector.

Potential limitations include:

  • Low-light environments
  • Poor camera quality
  • Strong video compression
  • Occluded faces
  • Extreme head poses
  • Multiple faces
  • Network-induced frame drops
  • Legitimate unusual facial motion
  • False positives from environmental changes
  • New deepfake generation techniques
  • Device-specific capture behavior

The risk score is an analytical signal and should be incorporated into an organization's broader fraud and compliance decision process.


πŸ” Privacy Considerations

Video KYC data can contain highly sensitive personal information.

A production implementation should consider:

  • Encryption in transit
  • Encryption at rest
  • Private object storage
  • Short media retention periods
  • Access-controlled evidence
  • Tenant isolation
  • Audit logs
  • Secure deletion policies
  • Least-privilege IAM
  • No unnecessary storage of raw biometric data

Do not expose uploaded KYC videos or evidence frames through public URLs.


πŸš€ Future Roadmap

Potential future improvements:

Detection

  • Dedicated deepfake classifier
  • Improved temporal transformer models
  • Audio deepfake detection
  • Stronger replay detection
  • Device/camera fingerprinting
  • Screen-reflection detection
  • Advanced face embeddings
  • Depth estimation
  • Multi-face attack handling

Infrastructure

  • Background processing workers
  • Redis-based session state
  • MongoDB Atlas integration
  • Amazon S3 integration
  • Horizontal scaling
  • GPU inference workers
  • Queue-based processing

Product

  • Tenant dashboard
  • Fraud analyst dashboard
  • Session history
  • Evidence explorer
  • Risk analytics
  • API keys
  • Webhooks
  • KYC provider integrations
  • Configurable risk thresholds
  • Organization-specific policies

🎯 Intended Use Cases

ShadowKYC can be used as an additional integrity layer for:

  • Fintech onboarding
  • Banking KYC
  • Insurance onboarding
  • Digital lending
  • Remote account opening
  • Government digital services
  • Telecom onboarding
  • High-value transaction verification
  • Remote employee verification
  • Fraud investigation workflows

🧩 Design Philosophy

ShadowKYC follows four core principles:

1. Multi-Signal Detection

No single visual signal is treated as sufficient evidence.

2. Temporal Analysis

Video should be analyzed across time, not only as independent images.

3. Cross-Layer Correlation

Contradictory signals can provide stronger evidence than isolated anomalies.

4. Explainable Risk

The system should provide reasons and evidence alongside the final risk score.


πŸ“Œ Important Terminology

Term Meaning
Liveness Evidence that the presented subject behaves like a live person
Replay Attack Presenting previously recorded KYC footage
Presentation Attack Presenting a photo, screen, mask, or other artifact as the subject
Face Swap Replacing one person's facial identity with another
Identity Drift Significant change in facial representation during a session
Temporal Consistency Whether facial behavior remains coherent over time
Texture Artifact Abnormal visual texture caused by synthesis, manipulation, or processing
Geometry Drift Abnormal change in facial landmark relationships
Lip-Sync Consistency between mouth movement and speech activity
Meta Fusion Combining multiple detection signals into one risk assessment

πŸ“œ Disclaimer

ShadowKYC is a security and fraud-analysis prototype intended to provide risk indicators and supporting evidence.

It should not be treated as a guaranteed deepfake detector or as the sole basis for identity, compliance, or financial decisions.

Production deployments should be evaluated against representative datasets, attack scenarios, regulatory requirements, privacy requirements, and organization-specific risk policies.


πŸ‘¨β€πŸ’» Development

This project is actively developed as a research and engineering project focused on:

Computer Vision
        +
Video Forensics
        +
Fraud Detection
        +
Real-Time Systems
        +
KYC Security

πŸ“„ License

Add the project's chosen license here before public production distribution.


⭐ Project Goal

Make Video KYC harder to fool by validating not only who is being presented, but whether the video itself appears trustworthy.

About

πŸ›‘οΈ Multi-layer KYC video integrity system for detecting deepfakes, replay attacks, face manipulation, and presentation attacks with explainable risk scoring.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages