An unsupervised machine learning approach to detect fraudulent credit card transactions using DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering algorithm — no labeled training data required.
This project implements an unsupervised anomaly detection system for identifying potentially fraudulent credit card transactions. Unlike supervised approaches, DBSCAN doesn't require labeled training data and can identify outliers (potential fraud) based on transaction patterns and density.
- No labeled data required - Works as an unsupervised approach
- Identifies outliers - Transactions marked as noise (-1) may indicate fraud
- Handles varying density - Can detect fraud patterns of different shapes
- Robust to noise - Doesn't force every point into a cluster
- 📊 Automated Data Processing - Handles missing values and infinite entries
- 🔄 Feature Standardization - Normalizes data for optimal clustering
- 🎯 DBSCAN Clustering - Density-based anomaly detection
- 📈 PCA Visualization - 2D projection of high-dimensional data
- 🧩 Silhouette Score - Quality metric for cluster separation
- 🎨 Plots - Visual representation of clusters and outliers
- Python 3.7 or higher
- Google Colab (recommended) or local Jupyter environment
pip install pandas numpy matplotlib scikit-learn
git clone https://github.com/zain-cs/Credit-Card-Fraud-DBSCAN.git cd Credit-Card-Fraud-DBSCAN
- Upload
DBSCAN_clustering.ipynbto Google Colab - Run all cells sequentially
- Upload your credit card transaction CSV when prompted
- View results and visualizations
- Open
DBSCAN_clustering.ipynbin Jupyter Notebook/Lab - Place your dataset CSV in the same directory (or update the file path in the notebook)
- Run all cells
Credit-Card-Fraud-DBSCAN/ │ ├── DBSCAN_clustering.ipynb # Main notebook: preprocessing, clustering, evaluation ├── README.md # This file └── LICENSE # MIT License
This project works with credit card transaction datasets containing numerical features. The expected format:
- Features: V1, V2, ..., V28 (PCA-transformed features)
- Time: Seconds elapsed between transactions
- Amount: Transaction amount
- Class (optional): 0 = Normal, 1 = Fraud (for validation only)
Kaggle Credit Card Fraud Detection Dataset
Time,V1,V2,V3,...,V28,Amount,Class 0,-1.359807134,-0.072781,...,149.62,0 1,-1.358354062,1.191857,...,2.69,0
- Remove infinite and missing values
- Select only numerical features
- Standardize features (mean=0, std=1)
DBSCAN(eps=1.5, min_samples=5)
- eps: Maximum distance between neighbors (tunable)
- min_samples: Minimum points to form a dense region
- Points labeled as
-1are outliers (potential fraud) - Dense clusters represent normal transaction patterns
- Noise points deviate significantly from normal behavior
- Silhouette Score: Measures cluster cohesion (excluding noise)
- Cluster Distribution: Shows normal vs. outlier counts
- PCA Visualization: 2D representation of clusters
📊 Cluster Distribution: 0 280000 -1 1500 1 500
- Cluster 0: Normal transactions (majority)
- Cluster -1: Outliers (potential fraud)
- Other clusters: Distinct transaction patterns
🔴 High Priority: Transactions in cluster -1 (noise/outliers)
🟡 Medium Priority: Small, isolated clusters
🟢 Low Priority: Large, dense clusters
Adjust DBSCAN parameters for your dataset:
DBSCAN(eps=0.5, min_samples=3)
DBSCAN(eps=2.0, min_samples=10)
Guidelines:
- Smaller eps → More outliers detected (higher sensitivity)
- Larger min_samples → Fewer small clusters (more conservative)
Contributions are welcome! Please follow these steps:
- Fork the repository
- Create a feature branch (
git checkout -b feature/AmazingFeature) - Commit your changes (
git commit -m 'Add some AmazingFeature') - Push to the branch (
git push origin feature/AmazingFeature) - Open a Pull Request
This project is licensed under the MIT License - see the LICENSE file for details.
Zain
GitHub: @zain-cs
- Scikit-learn for machine learning tools
- Kaggle for providing datasets
- Ester, M., et al. (1996). "A density-based algorithm for discovering clusters"
- Credit Card Fraud Detection Dataset (Kaggle)
- Scikit-learn DBSCAN Documentation
⭐ If you found this project helpful, please consider giving it a star!