This project predicts daily rainfall using an XGBoost regression model. It includes exploratory data analysis, data cleaning, time-series sequence generation, normalization, hyperparameter search, model evaluation, and prediction plotting.
For each prediction, the workflow uses a 1-day window of past weather variables and past rainfall values to predict the next day's 24-hour rainfall. The model does not use future information when building an input sequence.
The main focus is the rainfall prediction task. XGBoost is the modeling method, while preprocessing and sequence generation support the end-to-end prediction workflow.
This project demonstrates:
- End-to-end daily rainfall forecasting workflow, from raw weather data to model output.
- Multivariate time-series preprocessing using past weather variables and past rainfall values.
- Configurable lag feature construction for next-day rainfall prediction. The default run uses 1 previous day.
- XGBoost regression with TimeSeriesSplit cross-validation.
- Normalized MAAPE (%) as the main metric for rainfall data with many zero values.
- Supporting MAAPE, MAE, and RMSE metrics.
- Out-of-fold actual-vs-predicted CSV output and visualization for result inspection.
The notebook follows a compact analytical workflow aligned with common industry process models:
- CRISP-DM from IBM: data understanding, data preparation, modeling, evaluation, and deployment planning.
- SEMMA from SAS: sample, explore, modify, model, and assess.
- Modern ML lifecycle guidance from Microsoft/Azure Databricks: data preparation, feature processing, model training, evaluation, and reproducibility artifacts.
For this repository, raw data audit, preprocessing, validation, and EDA are kept as separate stages. The notebook first checks the raw dataset, fixes the identified data-quality issues, validates the cleaned result, and then analyzes the validated data before feature construction and modeling. The repository is intended as a reproducible forecasting workflow, not a production MLOps system.
| File | Description |
|---|---|
XGBoost_Rainfall_Prediction.ipynb |
Step-by-step notebook version of the workflow |
train_xgboost_rainfall.py |
Python script for preprocessing, training, evaluation, and output generation |
sample_weather_data.csv |
Synthetic sample dataset used by default for public demo/testing |
requirements.txt |
Python dependencies, including CuPy and its CUDA Toolkit components for optional NVIDIA/CUDA GPU mode |
LICENSE |
MIT License for the project code and included synthetic sample data |
The notebook is the main file to read and run. It is organized like a research workflow: configure the experiment, inspect the raw data, preprocess it, validate the cleaned dataset, run EDA, build lagged features, train and compare models, evaluate the result, and export the outputs.
The Python script acts as the backend for reusable functions. It keeps data preparation, feature generation, scaling, cross-validation, model evaluation, and output writing consistent without making the notebook too crowded. For normal experiments, change the configuration cell in the notebook first; edit the script only when changing the underlying workflow logic.
The original private run used weather observations provided with permission by BMKG Maritim Tanjung Perak in Surabaya City. The observations represent the BMKG Trunojoyo observation area in Sumenep Regency. The private raw dataset is not redistributed because access depends on local permission.
A synthetic sample dataset, sample_weather_data.csv, is included so this
project can be downloaded and run without access to private raw data. The sample
file follows the same column structure as the private/local dataset used during
development, but its weather values are synthetic.
By default, the notebook uses the included synthetic sample file:
data=Path("sample_weather_data.csv")To run another dataset with the same column structure, place the CSV file in the project folder and update the notebook configuration cell:
data=Path("your_weather_data.csv")The Python script also defaults to the included synthetic sample file. From the terminal, pass a file explicitly for another dataset:
python train_xgboost_rainfall.py --data your_weather_data.csvPrivate raw dataset filename patterns such as export-FKLIM-*.csv are
intentionally ignored by Git so private datasets are not committed accidentally.
This protection applies to Git commands. When using GitHub's browser uploader,
do not manually select a private export-FKLIM-*.csv file.
The expected dataset columns are listed below in the notebook display/model order:
| Column | Description |
|---|---|
DATA TIMESTAMP |
Observation date and time |
TEMPERATURE AVG C |
Average air temperature in Celsius |
SUNSHINE 24H H |
24-hour sunshine duration in hours |
QFF 24H MEAN MB |
24-hour mean sea-level pressure in millibars |
REL HUMIDITY AVG PC |
Average relative humidity in percent |
WIND SPEED 24H MEAN MS |
24-hour mean wind speed in meters per second |
EVAPORATION MM |
Evaporation in millimeters |
RAINFALL 24H MM |
24-hour rainfall in millimeters; also used as the prediction target |
For readability, notebook dataset tables display columns as date first, input
weather variables next, and the rainfall target RAINFALL 24H MM as the last
column. The same order is used in lagged feature tables: X weather variables
first, historical rainfall last.
The repository is set up as a runnable reference experiment, not as a locked
one-off report. The default run uses a sample dataset, a 1-day lag window,
TimeSeriesSplit(n_splits=10), and the included hyperparameter grid. The same
workflow can be reused for follow-up experiments by changing selected settings
in the notebook configuration cell.
The setup keeps the project reproducible while still leaving room to test reasonable alternatives.
Open the notebook and edit the configuration cell near the top:
args = Namespace(
data=Path("sample_weather_data.csv"),
date_col=DEFAULT_DATE_COL,
target_col=DEFAULT_TARGET_COL,
lag=1,
n_splits=10,
colsample_bytree=[0.6, 0.8, 1.0],
learning_rates=[0.05, 0.1, 0.3],
max_depths=[4, 6, 8],
n_estimators=[100, 300, 500],
subsamples=[0.6, 0.8, 1.0],
seed=42,
n_jobs=-1,
search_workers=-1,
device="cpu",
tree_method="exact",
base_score="train_mean",
reg_lambda=None,
reg_alpha=None,
gamma=None,
min_child_weight=None,
max_bin=None,
verbose=1,
output_dir=Path("outputs"),
keep_runs=1,
show_plot=False,
zero_codes=[8888.0, 9999.0],
include_target_history=True,
prepare_only=False,
)The configuration is separated into practical experiment controls:
| Group | Settings | Purpose |
|---|---|---|
| Data setup | data, date_col, target_col |
Select the CSV file and identify the time and rainfall target columns. |
| Forecast design | lag, n_splits |
Control the look-back window and TimeSeriesSplit fold count. |
| Grid-search hyperparameters | colsample_bytree, learning_rates, max_depths, n_estimators, subsamples |
Define the XGBoost regression settings tested across combinations. |
| Manual-check settings | tree_method, base_score, reg_lambda, reg_alpha, gamma, min_child_weight, max_bin |
Keep split construction, initial prediction, and regularization choices explicit when comparing with a manual calculation. |
| Reproducibility and compute | seed, n_jobs, search_workers, device |
Control repeatability, model threads, parallel grid combinations, and CPU/GPU selection. |
| Workflow and data rules | zero_codes, include_target_history, prepare_only |
Set special missing-value codes, include past rainfall among model inputs, or stop after data preparation. |
| Output controls | verbose, output_dir, keep_runs, show_plot |
Control progress messages, output location, retained runs, and interactive plot display. |
| Plot settings | Local plot_cfg = Namespace(...) or settings=Namespace(...) inside plotting cells |
Keep each plot's size, DPI, line width, and grid setting close to the plot it affects. |
The tuned XGBoost hyperparameters in this project are colsample_bytree, learning_rate, max_depth, n_estimators, and subsample. The notebook names learning_rates, max_depths, and subsamples are plural because each one stores the values tested by the grid search.
| Setting | Library or workflow meaning | Manual-check relevance |
|---|---|---|
tree_method |
XGBoost tree-construction method such as exact, approx, or hist. |
Directly affects how candidate splits are searched. Use exact for small manual split-gain checks. |
| Fixed objective | The workflow uses XGBoost regression with reg:squarederror. |
Defines the gradient and hessian behavior used by boosting, but it is not treated as a notebook experiment control. |
base_score |
Initial prediction before trees are added. | Can be None, a fixed value such as 0.5, or "train_mean" per fold. This is important for manual first-step calculations. |
learning_rates |
XGBoost learning_rate, also called eta. |
Shrinks each tree's contribution. |
max_depths |
XGBoost max_depth. |
Limits tree depth and therefore the number of split levels. |
n_estimators |
Number of boosting trees. | Controls how many sequential trees are added. |
subsamples |
XGBoost subsample. |
Row sampling ratio for each tree. |
colsample_bytree |
Column sampling ratio for each tree. | Controls how many features are available to a tree. |
gamma, min_child_weight, reg_alpha, reg_lambda, max_bin |
Optional split/regularization settings outside the grid. | Keep as None to use XGBoost defaults, or set explicit values to match a manual experiment. |
The notebook reference configuration uses CPU, tree_method="exact", and base_score="train_mean" because those settings are easier to compare with a small manual XGBoost calculation. If device="gpu", use tree_method="hist" because modern XGBoost GPU training does not support exact.
These parts define the current workflow. They can still be changed, but doing so changes the experiment design and usually belongs in the Python functions, not only in the notebook configuration cell:
- Model family:
XGBRegressor. - Main selection metric: mean CV Normalized MAAPE (%).
- Supporting metrics: MAAPE, MAE, and RMSE.
- Fold-based target scaling and prediction denormalization.
- Negative rainfall prediction clipping to
0. - Missing-value rules for this dataset.
- Min-Max normalization fitted inside each TimeSeriesSplit train fold.
- Pure TimeSeriesSplit cross-validation without a separate 80:20 split.
- Feature construction using numeric weather variables plus historical rainfall.
In short, the notebook configuration cell is meant for dataset, lag, TimeSeriesSplit, device, tree method, grid-search experiments, and manual-calculation settings. Plot appearance is kept inside each plotting cell. Deeper changes such as a different model family, normalization method, preprocessing rule, or metric formula belong in train_xgboost_rainfall.py.
The notebook and command-line defaults use device="cpu". During grid search,
search_workers=-1 assigns independent hyperparameter combinations to every
logical CPU thread. Available threads are divided across active fits, so broad
grids normally use one internal thread per fit while short or nearly completed
searches give more threads to each remaining fit. The selected-model and
default-model evaluations run sequentially with
n_jobs=-1, allowing each fit to use all available XGBoost threads.
Fold-local scaling, flattened lag arrays, and CPU/GPU model arrays are prepared
once and reused by every grid combination.
Tree-method rule:
| Device setting | Allowed tree method | Recommended use |
|---|---|---|
auto |
auto, exact, approx, hist |
Public/default CPU-compatible run |
cpu |
auto, exact, approx, hist |
exact for manual checks, hist for faster CPU training |
gpu |
hist only |
NVIDIA/CUDA GPU training and prediction |
Use device="cpu" or --device cpu to explicitly request CPU. Use
device="gpu" or --device gpu to request an NVIDIA/CUDA discrete GPU through
XGBoost's device="cuda" setting. Tree construction is controlled separately
with tree_method. When device is auto or cpu, tree_method can be
auto, exact, approx, or hist. This keeps exact available for small
manual-calculation checks. When device is gpu, tree_method is
hist; the workflow raises a clear error for GPU with exact, approx, or
auto instead of silently changing the setting. This matters because
tree_method affects split construction and therefore affects manual
calculation comparisons. The official XGBoost GPU documentation uses
XGBRegressor(tree_method="hist", device="cuda") with CuPy arrays for model
fit and predict. See the
official XGBoost GPU support documentation.
Integrated CPU GPUs are not targeted by this workflow. If GPU is requested but
nvidia-smi does not report an NVIDIA GPU, or CuPy cannot access CUDA, the
script raises a clear error instead of silently falling back to CPU.
GPU support depends on the local XGBoost installation, operating system,
driver, CUDA/runtime setup, and hardware. If GPU training fails, switch the
setting back to cpu or auto.
The single requirements file uses cupy-cuda12x[ctk] so CuPy has the CUDA
runtime and header components needed by GPU prediction. After pulling this
revision, rerun pip install -r requirements.txt before selecting GPU mode.
CPU and auto modes use NumPy CPU arrays for XGBoost fit and predict. GPU mode
uses CuPy GPU arrays for both XGBoost fit and predict, then converts predictions
back to NumPy only after prediction so metrics, CSV export, and plots can be
created with the standard pandas/NumPy workflow.
Use device="cpu" with tree_method="exact" for small manual-calculation
checks. Use device="cpu" with tree_method="hist" for faster CPU training.
Use device="gpu" with tree_method="hist" for NVIDIA/CUDA GPU runs.
Use search_workers=-1 for device-aware scheduling: it resolves to the
available logical CPU threads in CPU mode and one worker in GPU mode. A positive
value caps CPU combinations; GPU mode accepts -1 or 1. n_jobs=-1 is
retained for final and default evaluations, while parallel CPU grid fits divide
the available threads to avoid oversubscription.
GPU mode is optional, not an automatic speed guarantee. This project evaluates
many small fold models, so CUDA setup and per-model launch overhead can exceed
the saved training time. Model arrays already use float32 and the included
dataset fits comfortably in RAM. XGBoost external-memory or disk-backed input is
therefore not enabled; it is intended for data that does not fit in memory.
The notebook prints a compact compute runtime report. CPU mode shows the
processor and configured XGBoost runtime, while GPU mode also shows the detected
NVIDIA GPU index, name, driver, and total GPU memory when nvidia-smi is
available.
The main GPU verification is shown by the booster_devices_after_fit and
model_array_backends columns in the result tables and output CSV files. A
value such as cuda:0 means the fitted XGBoost booster used CUDA. A value such
as cupy_gpu_arrays means the model fit/predict arrays were on GPU. CPU mode
shows cpu and numpy_cpu_arrays.
Known non-critical XGBoost device-mismatch warning noise is handled in code. Important warnings such as invalid parameters, incompatible packages, or failed GPU training are still allowed to appear.
- Define the workflow configuration and selected dataset.
- Load the selected weather dataset.
- Audit the raw dataset to inspect size, date range, missing values, special codes, duplicate dates, and rainfall balance.
- Apply preprocessing rules: chronological sorting, duplicate-date handling, X-variable interpolation, and rainfall target filling.
- Validate the cleaned dataset before feature construction.
- Run EDA on the validated dataset to inspect rainfall behavior, input-variable behavior, seasonal patterns, extreme events, and temporal signals.
- Build sliding-window samples using the configured lag value.
- Evaluate models with the configured
TimeSeriesSplitfold count. - For each fold, normalize input features using Min-Max values fitted from that fold's train data only.
- For each fold, scale the target during training and denormalize predictions back to millimeters.
- Clip negative rainfall predictions to
0after denormalization. - Tune XGBoost models with TimeSeriesSplit cross-validation and grid search.
- Compare the selected model with a default XGBoost model using the same CV folds.
- Select the best hyperparameter combination using the lowest mean CV Normalized MAAPE (%).
- Save out-of-fold predictions, metrics, plots, and run metadata.
flowchart LR
A["CSV weather data"] --> B["Raw data audit"]
B --> C["Data preprocessing"]
C --> D["Cleaned data validation"]
D --> E["Exploratory data analysis"]
E --> F["Lagged feature construction"]
F --> G["TimeSeriesSplit folds"]
G --> H["Fold-based normalization"]
H --> I["XGBoost grid search"]
I --> J["Default model comparison"]
J --> K["Metrics, out-of-fold CSV outputs, and plot"]
The notebook separates raw-data checks, cleaned-data validation, and EDA so each stage has a clear purpose.
The raw data audit checks the selected CSV before values are changed. It shows:
- Raw dataset overview: row count, column count, date range, duplicate dates, and number of numeric columns.
- Full raw dataset table using source-like decimal formatting, with the rainfall target displayed as the last column.
- Missing-value counts and percentages.
- Counts of special values such as
8888and9999. - Rainfall balance: missing rainfall days, zero-rainfall days, positive-rainfall days, valid positive rainfall days, and valid rainfall summary values.
Preprocessing then handles the issues found in the raw check: special codes in
input variables are treated as missing and interpolated, special codes in
rainfall are filled with 0, missing rainfall is filled with 0, duplicate
dates are removed, and the data is sorted chronologically.
Cleaned data validation verifies the modeling dataset before lag construction. It shows remaining missing values, remaining special codes, duplicate dates, expected daily date gaps, rainfall range, zero/positive rainfall counts, numeric summary, and the prepared dataset table.
EDA is then performed on the validated dataset:
- Rainfall time-series plot. Spikes indicate heavier rainfall days, while flat zero periods indicate no recorded rainfall.
- Combined time-series subplot for input weather variables. The rainfall target is excluded from this subplot because it already has its own rainfall time-series plot.
- Rainfall distribution plot. This shows whether the data is dominated by zero/low-rainfall days or contains many high-rainfall events.
- Feature correlation matrix and heatmap. These show linear relationships between variables; they do not prove causation.
- Calendar-month rainfall distribution table and boxplot. These show whether rainfall behavior differs by month, which helps inspect seasonal rainfall patterns.
- Extreme rainfall summary and event table. These identify rare heavy-rainfall days that may be harder for a forecasting model to predict.
- Rainfall autocorrelation by lag. This checks whether past rainfall contains temporal signal for future rainfall.
- Lagged feature-to-target correlation. This checks whether each lagged weather input has a direct linear relationship with the target rainfall day.
- TimeSeriesSplit rainfall distribution comparison to review whether some test folds are wetter, drier, or more extreme than others.
These EDA outputs are used to understand the dataset and to identify sensible follow-up experiments. They do not automatically determine the final lag, cross-validation setup or hyperparameter grid. For example, a higher autocorrelation at another lag is a useful signal to test that lag, but the final choice still needs to be evaluated through model training and cross-validation.
This order prevents raw special codes such as 8888 or 9999 from distorting
the plots.
The preprocessing rules follow the handling notes for this dataset:
- Missing calendar dates are inserted so every lag step represents one day. The inserted X values follow the same interpolation rule, while the inserted rainfall target follows the same zero-fill rule.
8888and9999in input weather variables are treated as unavailable values, converted to missing values, and filled using linear interpolation.- Missing values in input weather variables are also filled using linear interpolation.
- If an input weather column still cannot be filled after interpolation because
it has no valid values, the workflow raises an error instead of filling X
with
0. 8888,9999, and missing values in the rainfall target are filled with0, because these rainfall entries are handled as no recorded rainfall for this dataset.- After filling missing values, weather columns are rounded to match the decimal precision rules defined in the training script.
The target variable is:
RAINFALL 24H MM
The default run uses a 1-day lag window. For each sample, the configured number of previous days is used to predict rainfall on the next day.
With the default 1-day lag and the expected weather data structure, the sequence input shape before flattening is:
X = (samples, 1, 7)
y = (samples, 1)
XGBoost receives each normalized lag window as a flattened tabular input during training.
The seven input features are:
TEMPERATURE AVG C
SUNSHINE 24H H
QFF 24H MEAN MB
REL HUMIDITY AVG PC
WIND SPEED 24H MEAN MS
EVAPORATION MM
RAINFALL 24H MM
The lagged feature order follows the same readable dataset order: X weather variables first and historical rainfall last.
The notebook displays the complete supervised lagged dataset after this step,
and each completed script/notebook run saves the same table as
lagged_dataset.csv.
Custom datasets use the same column names so the preprocessing and model workflow can run without code changes.
The project tunes five XGBoost hyperparameters:
colsample_bytree = [0.6, 0.8, 1.0]
learning_rates = [0.05, 0.1, 0.3]
max_depths = [4, 6, 8]
n_estimators = [100, 300, 500]
subsamples = [0.6, 0.8, 1.0]TimeSeriesSplit is fixed at 10 folds for every run and is not part of the
grid search.
Total combinations:
3 x 3 x 3 x 3 x 3 = 243
With 10 TimeSeriesSplit folds, the grid search performs:
243 x 10 = 2430 cross-validation fits
The XGBoost model used for the tuned search is:
XGBRegressor
The default comparison model uses XGBRegressor defaults for the tuned
hyperparameters, with the project seed passed internally as random_state for
reproducibility, while keeping the same manual-calculation settings from
the configuration cell.
These settings affect the equations used in a manual XGBoost calculation. The objective is fixed by the workflow, while the remaining settings are available in the notebook configuration cell:
| Setting | Role in manual calculation |
|---|---|
Fixed objective: reg:squarederror |
Defines the gradient and hessian formula for this regression workflow |
tree_method |
Controls how candidate splits are searched (exact, hist, etc.) |
base_score |
Initial prediction before the first tree |
reg_lambda |
L2 regularization term in leaf-weight calculation |
reg_alpha |
L1 regularization term in leaf-weight calculation |
gamma |
Minimum split gain required to create a split |
min_child_weight |
Minimum child-node hessian/weight requirement |
max_bin |
Number of histogram bins when using hist or approx |
They are not part of the grid search. They are fixed for a run and applied to every grid-search combination and to the default comparison model. The tuned grid remains:
colsample_bytree, learning_rate, max_depth, n_estimators, subsample
base_score can be configured in three ways:
| Setting | Meaning |
|---|---|
base_score=None |
Leave the initial prediction to XGBoost's default/automatic handling |
base_score=0.5 |
Use a fixed initial prediction of 0.5 in every TimeSeriesSplit fold |
base_score="train_mean" |
Use mean(y_train_scaled) separately inside each fold |
For manual-calculation checks, the notebook reference configuration uses
base_score="train_mean" because the calculation starts from the
training-target mean. The mean is calculated only from each fold's normalized
training target, so the fold test period is not used when setting the initial
prediction. Set base_score=None for XGBoost's default/automatic handling.
For a small manual-check setup, reduce the grid to one value per tuned parameter, use CPU, and choose a compact tree:
device = "cpu"
tree_method = "exact"
n_estimators = [1]
max_depths = [1]
learning_rates = [1.0]
subsamples = [1.0]
colsample_bytree = [1.0]
base_score = "train_mean"
reg_lambda = 1.0
reg_alpha = 0.0
gamma = 0.0
min_child_weight = 1.0The main evaluation metric is Normalized MAAPE (%):
MAAPE = mean(arctan2(abs(y_true - y_pred), abs(y_true)))
Normalized MAAPE = MAAPE / (pi / 2)
Normalized MAAPE (%) = Normalized MAAPE * 100
MAAPE is used because rainfall data often contains many zero values, where ordinary MAPE can become unstable. The grid search uses mean CV Normalized MAAPE (%) for model selection.
arctan2 is used instead of arctan(abs(error) / abs(actual)) because it
avoids manual division by the actual rainfall value. This keeps the metric
defined when the actual rainfall is 0. When both actual and predicted rainfall
are 0, the MAAPE contribution is 0; when the actual rainfall is 0 but the
prediction is not 0, the contribution becomes the maximum arctangent penalty.
The normalization step converts the original MAAPE angle from the range
0 to pi/2 into a more readable 0% to 100% scale. The original MAAPE
angle is also reported as a supporting metric, while MAE and RMSE are reported
in millimeters. Metric tables are displayed with 4 decimal places in the
notebook.
- Lower Normalized MAAPE (%), MAAPE, MAE, and RMSE values indicate better predictive performance.
- The best tuned model is selected using mean CV Normalized MAAPE (%), not MAE or RMSE. Plain MAAPE is displayed only as an additional metric.
- MAE and RMSE are reported in millimeters, so they are easier to interpret in the original rainfall unit.
- A tuned model can improve Normalized MAAPE (%) while not improving every supporting metric. This can happen because each metric penalizes errors differently.
- The final prediction CSV contains out-of-fold predictions. Each row is predicted by a model that was trained only on data before that fold's test period.
- Results produced with
sample_weather_data.csvare intended for reproducible public demonstration. They are not interpreted as the performance of private raw data.
Rainfall is a non-negative variable, so negative model predictions are clipped
to 0 after predictions are denormalized back to millimeters. This clipping is
applied before evaluation/scoring, CSV export, and plotting.
- The dataset represents one observation area, so the result is not assumed to generalize directly to other regions without retraining or further testing.
- Rainfall data contains many zero-rainfall days and sudden high-rainfall events. The model may follow the general pattern better than sharp extreme peaks.
- The default run uses a 1-day lag window. Other lag lengths may produce different results and can be tested from the notebook configuration cell.
- The experiment uses non-nested TimeSeriesSplit cross-validation. The reported CV score is useful for comparing hyperparameter combinations, but it may still be slightly optimistic because the same CV process is used for tuning and reporting.
- The earliest initial training segment does not have out-of-fold predictions because TimeSeriesSplit preserves chronological order.
- The model only uses the variables available in the CSV file. Additional predictors such as radar, satellite, numerical weather prediction, or broader climate indices may improve performance if available.
- Exact results can vary slightly across package versions and runtime environments, although the code sets a random seed for reproducibility.
- Linear interpolation is applied as an offline data-cleaning rule and can use observations on both sides of a gap. A real-time operational pipeline should replace it with a causal imputation rule that only uses information available before the forecast timestamp.
Python 3.12 is recommended. The code requires Python 3.10 or newer.
pip install -r requirements.txtThe same requirements file is used for CPU and optional NVIDIA/CUDA GPU runs.
CPU users can keep device="cpu" or device="auto" in the notebook. GPU users
use device="gpu" with tree_method="hist".
Run the notebook:
XGBoost_Rainfall_Prediction.ipynb
Check preprocessing without training:
python train_xgboost_rainfall.py --prepare-onlyRun full training:
python train_xgboost_rainfall.pyForce CPU:
python train_xgboost_rainfall.py --device cpuRequest GPU:
python train_xgboost_rainfall.py --device gpuShow the final plot after training:
python train_xgboost_rainfall.py --show-plotRun a smaller grid search:
python train_xgboost_rainfall.py --colsample-bytree 0.8 --learning-rates 0.1 --max-depths 4 --n-estimators 100 --subsamples 0.8- The notebook and script default to the included synthetic sample dataset, so the public workflow can run without private data.
- A random seed is set for reproducibility, but exact results can still vary slightly across package versions and runtime environments.
- Full XGBoost search evaluates 243 hyperparameter combinations across 10 TimeSeriesSplit folds. Train-fold metrics and detailed out-of-fold predictions are calculated only for the selected combination, avoiding unnecessary train inference for every candidate. The selected and untuned comparison setups are then evaluated once on the same folds for their detailed outputs.
- Use
--prepare-onlyto verify data loading and preprocessing without model training. - Use the smaller grid-search command above for a faster functional check.
Each completed run creates an output folder under outputs/. By default, only
the latest completed run folder is kept.
Main output files:
| File | Description |
|---|---|
lagged_dataset.csv |
Complete supervised lagged dataset used for modeling |
hyperparameter_results.csv |
All grid-search combinations sorted by mean_cv_normalized_maape_percent |
best_cv_fold_metrics.csv |
Fold-by-fold metrics for the best grid-search model |
best_cv_predictions.csv |
Out-of-fold actual vs predicted rainfall for the best model |
best_cv_prediction_plot.png |
Actual vs predicted out-of-fold plot for the best model |
default_cv_fold_metrics.csv |
Fold-by-fold metrics from the default model |
default_cv_predictions.csv |
Out-of-fold actual vs predicted rainfall from the default model |
default_cv_prediction_plot.png |
Default model out-of-fold plot |
best_vs_default_comparison.csv |
Best grid-search model vs default model |
run_metadata.json |
Configuration, compute runtime, feature columns, split information, and preprocessing details |
The notebook also displays a compact final summary:
Best model Normalized MAAPE (%):
Best model MAAPE:
Best model MAE:
Best model RMSE:
Default model vs tuned model:
Generated outputs, model artifacts, Python caches, and notebook checkpoints are ignored by Git.
This project is released under the MIT License. The license applies to the project code and the included synthetic sample dataset. Private raw datasets are not redistributed in this repository and are not covered as public datasets by this license.