A Machine Learning based web application that predicts customer churn probability using customer subscription and behavioral data.
The project focuses on identifying high-risk customers early using classification techniques and recall optimization.
- Predicts customer churn probability
- Logistic Regression based ML model
- Recall focused threshold optimization
- Interactive Streamlit dashboard
- Customer risk classification
- Churn risk visualization
- Retention recommendation system
- Data Cleaning
- Exploratory Data Analysis
- Feature Encoding
- Feature Scaling
- Logistic Regression Training
- Threshold Optimization
- Model Deployment using Streamlit
IBM Telco Customer Churn Dataset
Dataset contains:
- 7043 customer records
- Demographic information
- Subscription details
- Billing information
- Customer churn labels
| Metric | Score |
|---|---|
| Accuracy | 71% |
| Recall | 81% |
Recall optimization is applied to reduce false negatives and improve detection of customers likely to churn.
- Python
- Pandas
- NumPy
- Scikit-learn
- Streamlit
- Plotly
- Joblib
Customer-Churn-Prediction-System
├── data
│ └── Telco-Customer-Churn.csv
│
├── models
│ ├── churn_model.pkl
│ ├── scaler.pkl
│ └── features.pkl
│
├── app.py
├── train_model.py
├── preprocessing.py
├── requirements.txt
└── README.md
Clone repository
git clone <repo-link>Install dependencies
pip install -r requirements.txtTrain model
python train_model.pyStart Streamlit App
streamlit run app.py- Add advanced ML models
- Add SHAP explainability
- Deploy using Streamlit Cloud