Skip to content

Repository files navigation

Rainfall Forecasting with XGBoost and Grid Search

Python 3.10+ XGBoost Regression Jupyter Notebook License: MIT

This project predicts daily rainfall using an XGBoost regression model. It includes exploratory data analysis, data cleaning, time-series sequence generation, normalization, hyperparameter search, model evaluation, and prediction plotting.

For each prediction, the workflow uses a 1-day window of past weather variables and past rainfall values to predict the next day's 24-hour rainfall. The model does not use future information when building an input sequence.

The main focus is the rainfall prediction task. XGBoost is the modeling method, while preprocessing and sequence generation support the end-to-end prediction workflow.

Project Highlights

This project demonstrates:

  • End-to-end daily rainfall forecasting workflow, from raw weather data to model output.
  • Multivariate time-series preprocessing using past weather variables and past rainfall values.
  • Configurable lag feature construction for next-day rainfall prediction. The default run uses 1 previous day.
  • XGBoost regression with TimeSeriesSplit cross-validation.
  • Normalized MAAPE (%) as the main metric for rainfall data with many zero values.
  • Supporting MAAPE, MAE, and RMSE metrics.
  • Out-of-fold actual-vs-predicted CSV output and visualization for result inspection.

Workflow Alignment

The notebook follows a compact analytical workflow aligned with common industry process models:

For this repository, raw data audit, preprocessing, validation, and EDA are kept as separate stages. The notebook first checks the raw dataset, fixes the identified data-quality issues, validates the cleaned result, and then analyzes the validated data before feature construction and modeling. The repository is intended as a reproducible forecasting workflow, not a production MLOps system.

Files

File Description
XGBoost_Rainfall_Prediction.ipynb Step-by-step notebook version of the workflow
train_xgboost_rainfall.py Python script for preprocessing, training, evaluation, and output generation
sample_weather_data.csv Synthetic sample dataset used by default for public demo/testing
requirements.txt Python dependencies, including CuPy and its CUDA Toolkit components for optional NVIDIA/CUDA GPU mode
LICENSE MIT License for the project code and included synthetic sample data

Notebook and Python Script Roles

The notebook is the main file to read and run. It is organized like a research workflow: configure the experiment, inspect the raw data, preprocess it, validate the cleaned dataset, run EDA, build lagged features, train and compare models, evaluate the result, and export the outputs.

The Python script acts as the backend for reusable functions. It keeps data preparation, feature generation, scaling, cross-validation, model evaluation, and output writing consistent without making the notebook too crowded. For normal experiments, change the configuration cell in the notebook first; edit the script only when changing the underlying workflow logic.

Data Availability

The original private run used weather observations provided with permission by BMKG Maritim Tanjung Perak in Surabaya City. The observations represent the BMKG Trunojoyo observation area in Sumenep Regency. The private raw dataset is not redistributed because access depends on local permission.

A synthetic sample dataset, sample_weather_data.csv, is included so this project can be downloaded and run without access to private raw data. The sample file follows the same column structure as the private/local dataset used during development, but its weather values are synthetic.

By default, the notebook uses the included synthetic sample file:

data=Path("sample_weather_data.csv")

To run another dataset with the same column structure, place the CSV file in the project folder and update the notebook configuration cell:

data=Path("your_weather_data.csv")

The Python script also defaults to the included synthetic sample file. From the terminal, pass a file explicitly for another dataset:

python train_xgboost_rainfall.py --data your_weather_data.csv

Private raw dataset filename patterns such as export-FKLIM-*.csv are intentionally ignored by Git so private datasets are not committed accidentally. This protection applies to Git commands. When using GitHub's browser uploader, do not manually select a private export-FKLIM-*.csv file.

Data Dictionary

The expected dataset columns are listed below in the notebook display/model order:

Column Description
DATA TIMESTAMP Observation date and time
TEMPERATURE AVG C Average air temperature in Celsius
SUNSHINE 24H H 24-hour sunshine duration in hours
QFF 24H MEAN MB 24-hour mean sea-level pressure in millibars
REL HUMIDITY AVG PC Average relative humidity in percent
WIND SPEED 24H MEAN MS 24-hour mean wind speed in meters per second
EVAPORATION MM Evaporation in millimeters
RAINFALL 24H MM 24-hour rainfall in millimeters; also used as the prediction target

For readability, notebook dataset tables display columns as date first, input weather variables next, and the rainfall target RAINFALL 24H MM as the last column. The same order is used in lagged feature tables: X weather variables first, historical rainfall last.

Experiment Scope

The repository is set up as a runnable reference experiment, not as a locked one-off report. The default run uses a sample dataset, a 1-day lag window, TimeSeriesSplit(n_splits=10), and the included hyperparameter grid. The same workflow can be reused for follow-up experiments by changing selected settings in the notebook configuration cell.

The setup keeps the project reproducible while still leaving room to test reasonable alternatives.

Configurable from the Notebook

Open the notebook and edit the configuration cell near the top:

args = Namespace(
    data=Path("sample_weather_data.csv"),
    date_col=DEFAULT_DATE_COL,
    target_col=DEFAULT_TARGET_COL,
    lag=1,
    n_splits=10,
    colsample_bytree=[0.6, 0.8, 1.0],
    learning_rates=[0.05, 0.1, 0.3],
    max_depths=[4, 6, 8],
    n_estimators=[100, 300, 500],
    subsamples=[0.6, 0.8, 1.0],
    seed=42,
    n_jobs=-1,
    search_workers=-1,
    device="cpu",
    tree_method="exact",
    base_score="train_mean",
    reg_lambda=None,
    reg_alpha=None,
    gamma=None,
    min_child_weight=None,
    max_bin=None,
    verbose=1,
    output_dir=Path("outputs"),
    keep_runs=1,
    show_plot=False,
    zero_codes=[8888.0, 9999.0],
    include_target_history=True,
    prepare_only=False,
)

The configuration is separated into practical experiment controls:

Group Settings Purpose
Data setup data, date_col, target_col Select the CSV file and identify the time and rainfall target columns.
Forecast design lag, n_splits Control the look-back window and TimeSeriesSplit fold count.
Grid-search hyperparameters colsample_bytree, learning_rates, max_depths, n_estimators, subsamples Define the XGBoost regression settings tested across combinations.
Manual-check settings tree_method, base_score, reg_lambda, reg_alpha, gamma, min_child_weight, max_bin Keep split construction, initial prediction, and regularization choices explicit when comparing with a manual calculation.
Reproducibility and compute seed, n_jobs, search_workers, device Control repeatability, model threads, parallel grid combinations, and CPU/GPU selection.
Workflow and data rules zero_codes, include_target_history, prepare_only Set special missing-value codes, include past rainfall among model inputs, or stop after data preparation.
Output controls verbose, output_dir, keep_runs, show_plot Control progress messages, output location, retained runs, and interactive plot display.
Plot settings Local plot_cfg = Namespace(...) or settings=Namespace(...) inside plotting cells Keep each plot's size, DPI, line width, and grid setting close to the plot it affects.

XGBoost Hyperparameters and Manual-Check Notes

The tuned XGBoost hyperparameters in this project are colsample_bytree, learning_rate, max_depth, n_estimators, and subsample. The notebook names learning_rates, max_depths, and subsamples are plural because each one stores the values tested by the grid search.

Setting Library or workflow meaning Manual-check relevance
tree_method XGBoost tree-construction method such as exact, approx, or hist. Directly affects how candidate splits are searched. Use exact for small manual split-gain checks.
Fixed objective The workflow uses XGBoost regression with reg:squarederror. Defines the gradient and hessian behavior used by boosting, but it is not treated as a notebook experiment control.
base_score Initial prediction before trees are added. Can be None, a fixed value such as 0.5, or "train_mean" per fold. This is important for manual first-step calculations.
learning_rates XGBoost learning_rate, also called eta. Shrinks each tree's contribution.
max_depths XGBoost max_depth. Limits tree depth and therefore the number of split levels.
n_estimators Number of boosting trees. Controls how many sequential trees are added.
subsamples XGBoost subsample. Row sampling ratio for each tree.
colsample_bytree Column sampling ratio for each tree. Controls how many features are available to a tree.
gamma, min_child_weight, reg_alpha, reg_lambda, max_bin Optional split/regularization settings outside the grid. Keep as None to use XGBoost defaults, or set explicit values to match a manual experiment.

The notebook reference configuration uses CPU, tree_method="exact", and base_score="train_mean" because those settings are easier to compare with a small manual XGBoost calculation. If device="gpu", use tree_method="hist" because modern XGBoost GPU training does not support exact.

Core Workflow Choices

These parts define the current workflow. They can still be changed, but doing so changes the experiment design and usually belongs in the Python functions, not only in the notebook configuration cell:

  • Model family: XGBRegressor.
  • Main selection metric: mean CV Normalized MAAPE (%).
  • Supporting metrics: MAAPE, MAE, and RMSE.
  • Fold-based target scaling and prediction denormalization.
  • Negative rainfall prediction clipping to 0.
  • Missing-value rules for this dataset.
  • Min-Max normalization fitted inside each TimeSeriesSplit train fold.
  • Pure TimeSeriesSplit cross-validation without a separate 80:20 split.
  • Feature construction using numeric weather variables plus historical rainfall.

In short, the notebook configuration cell is meant for dataset, lag, TimeSeriesSplit, device, tree method, grid-search experiments, and manual-calculation settings. Plot appearance is kept inside each plotting cell. Deeper changes such as a different model family, normalization method, preprocessing rule, or metric formula belong in train_xgboost_rainfall.py.

Compute Device

The notebook and command-line defaults use device="cpu". During grid search, search_workers=-1 assigns independent hyperparameter combinations to every logical CPU thread. Available threads are divided across active fits, so broad grids normally use one internal thread per fit while short or nearly completed searches give more threads to each remaining fit. The selected-model and default-model evaluations run sequentially with n_jobs=-1, allowing each fit to use all available XGBoost threads. Fold-local scaling, flattened lag arrays, and CPU/GPU model arrays are prepared once and reused by every grid combination.

Tree-method rule:

Device setting Allowed tree method Recommended use
auto auto, exact, approx, hist Public/default CPU-compatible run
cpu auto, exact, approx, hist exact for manual checks, hist for faster CPU training
gpu hist only NVIDIA/CUDA GPU training and prediction

Use device="cpu" or --device cpu to explicitly request CPU. Use device="gpu" or --device gpu to request an NVIDIA/CUDA discrete GPU through XGBoost's device="cuda" setting. Tree construction is controlled separately with tree_method. When device is auto or cpu, tree_method can be auto, exact, approx, or hist. This keeps exact available for small manual-calculation checks. When device is gpu, tree_method is hist; the workflow raises a clear error for GPU with exact, approx, or auto instead of silently changing the setting. This matters because tree_method affects split construction and therefore affects manual calculation comparisons. The official XGBoost GPU documentation uses XGBRegressor(tree_method="hist", device="cuda") with CuPy arrays for model fit and predict. See the official XGBoost GPU support documentation. Integrated CPU GPUs are not targeted by this workflow. If GPU is requested but nvidia-smi does not report an NVIDIA GPU, or CuPy cannot access CUDA, the script raises a clear error instead of silently falling back to CPU.

GPU support depends on the local XGBoost installation, operating system, driver, CUDA/runtime setup, and hardware. If GPU training fails, switch the setting back to cpu or auto.

The single requirements file uses cupy-cuda12x[ctk] so CuPy has the CUDA runtime and header components needed by GPU prediction. After pulling this revision, rerun pip install -r requirements.txt before selecting GPU mode.

CPU and auto modes use NumPy CPU arrays for XGBoost fit and predict. GPU mode uses CuPy GPU arrays for both XGBoost fit and predict, then converts predictions back to NumPy only after prediction so metrics, CSV export, and plots can be created with the standard pandas/NumPy workflow.

Use device="cpu" with tree_method="exact" for small manual-calculation checks. Use device="cpu" with tree_method="hist" for faster CPU training. Use device="gpu" with tree_method="hist" for NVIDIA/CUDA GPU runs.

Use search_workers=-1 for device-aware scheduling: it resolves to the available logical CPU threads in CPU mode and one worker in GPU mode. A positive value caps CPU combinations; GPU mode accepts -1 or 1. n_jobs=-1 is retained for final and default evaluations, while parallel CPU grid fits divide the available threads to avoid oversubscription.

GPU mode is optional, not an automatic speed guarantee. This project evaluates many small fold models, so CUDA setup and per-model launch overhead can exceed the saved training time. Model arrays already use float32 and the included dataset fits comfortably in RAM. XGBoost external-memory or disk-backed input is therefore not enabled; it is intended for data that does not fit in memory.

The notebook prints a compact compute runtime report. CPU mode shows the processor and configured XGBoost runtime, while GPU mode also shows the detected NVIDIA GPU index, name, driver, and total GPU memory when nvidia-smi is available. The main GPU verification is shown by the booster_devices_after_fit and model_array_backends columns in the result tables and output CSV files. A value such as cuda:0 means the fitted XGBoost booster used CUDA. A value such as cupy_gpu_arrays means the model fit/predict arrays were on GPU. CPU mode shows cpu and numpy_cpu_arrays.

Known non-critical XGBoost device-mismatch warning noise is handled in code. Important warnings such as invalid parameters, incompatible packages, or failed GPU training are still allowed to appear.

Workflow

  1. Define the workflow configuration and selected dataset.
  2. Load the selected weather dataset.
  3. Audit the raw dataset to inspect size, date range, missing values, special codes, duplicate dates, and rainfall balance.
  4. Apply preprocessing rules: chronological sorting, duplicate-date handling, X-variable interpolation, and rainfall target filling.
  5. Validate the cleaned dataset before feature construction.
  6. Run EDA on the validated dataset to inspect rainfall behavior, input-variable behavior, seasonal patterns, extreme events, and temporal signals.
  7. Build sliding-window samples using the configured lag value.
  8. Evaluate models with the configured TimeSeriesSplit fold count.
  9. For each fold, normalize input features using Min-Max values fitted from that fold's train data only.
  10. For each fold, scale the target during training and denormalize predictions back to millimeters.
  11. Clip negative rainfall predictions to 0 after denormalization.
  12. Tune XGBoost models with TimeSeriesSplit cross-validation and grid search.
  13. Compare the selected model with a default XGBoost model using the same CV folds.
  14. Select the best hyperparameter combination using the lowest mean CV Normalized MAAPE (%).
  15. Save out-of-fold predictions, metrics, plots, and run metadata.
flowchart LR
    A["CSV weather data"] --> B["Raw data audit"]
    B --> C["Data preprocessing"]
    C --> D["Cleaned data validation"]
    D --> E["Exploratory data analysis"]
    E --> F["Lagged feature construction"]
    F --> G["TimeSeriesSplit folds"]
    G --> H["Fold-based normalization"]
    H --> I["XGBoost grid search"]
    I --> J["Default model comparison"]
    J --> K["Metrics, out-of-fold CSV outputs, and plot"]
Loading

Exploratory Data Analysis

The notebook separates raw-data checks, cleaned-data validation, and EDA so each stage has a clear purpose.

The raw data audit checks the selected CSV before values are changed. It shows:

  • Raw dataset overview: row count, column count, date range, duplicate dates, and number of numeric columns.
  • Full raw dataset table using source-like decimal formatting, with the rainfall target displayed as the last column.
  • Missing-value counts and percentages.
  • Counts of special values such as 8888 and 9999.
  • Rainfall balance: missing rainfall days, zero-rainfall days, positive-rainfall days, valid positive rainfall days, and valid rainfall summary values.

Preprocessing then handles the issues found in the raw check: special codes in input variables are treated as missing and interpolated, special codes in rainfall are filled with 0, missing rainfall is filled with 0, duplicate dates are removed, and the data is sorted chronologically.

Cleaned data validation verifies the modeling dataset before lag construction. It shows remaining missing values, remaining special codes, duplicate dates, expected daily date gaps, rainfall range, zero/positive rainfall counts, numeric summary, and the prepared dataset table.

EDA is then performed on the validated dataset:

  • Rainfall time-series plot. Spikes indicate heavier rainfall days, while flat zero periods indicate no recorded rainfall.
  • Combined time-series subplot for input weather variables. The rainfall target is excluded from this subplot because it already has its own rainfall time-series plot.
  • Rainfall distribution plot. This shows whether the data is dominated by zero/low-rainfall days or contains many high-rainfall events.
  • Feature correlation matrix and heatmap. These show linear relationships between variables; they do not prove causation.
  • Calendar-month rainfall distribution table and boxplot. These show whether rainfall behavior differs by month, which helps inspect seasonal rainfall patterns.
  • Extreme rainfall summary and event table. These identify rare heavy-rainfall days that may be harder for a forecasting model to predict.
  • Rainfall autocorrelation by lag. This checks whether past rainfall contains temporal signal for future rainfall.
  • Lagged feature-to-target correlation. This checks whether each lagged weather input has a direct linear relationship with the target rainfall day.
  • TimeSeriesSplit rainfall distribution comparison to review whether some test folds are wetter, drier, or more extreme than others.

These EDA outputs are used to understand the dataset and to identify sensible follow-up experiments. They do not automatically determine the final lag, cross-validation setup or hyperparameter grid. For example, a higher autocorrelation at another lag is a useful signal to test that lag, but the final choice still needs to be evaluated through model training and cross-validation.

This order prevents raw special codes such as 8888 or 9999 from distorting the plots.

Preprocessing Notes

The preprocessing rules follow the handling notes for this dataset:

  • Missing calendar dates are inserted so every lag step represents one day. The inserted X values follow the same interpolation rule, while the inserted rainfall target follows the same zero-fill rule.
  • 8888 and 9999 in input weather variables are treated as unavailable values, converted to missing values, and filled using linear interpolation.
  • Missing values in input weather variables are also filled using linear interpolation.
  • If an input weather column still cannot be filled after interpolation because it has no valid values, the workflow raises an error instead of filling X with 0.
  • 8888, 9999, and missing values in the rainfall target are filled with 0, because these rainfall entries are handled as no recorded rainfall for this dataset.
  • After filling missing values, weather columns are rounded to match the decimal precision rules defined in the training script.

Model Input

The target variable is:

RAINFALL 24H MM

The default run uses a 1-day lag window. For each sample, the configured number of previous days is used to predict rainfall on the next day.

With the default 1-day lag and the expected weather data structure, the sequence input shape before flattening is:

X = (samples, 1, 7)
y = (samples, 1)

XGBoost receives each normalized lag window as a flattened tabular input during training.

The seven input features are:

TEMPERATURE AVG C
SUNSHINE 24H H
QFF 24H MEAN MB
REL HUMIDITY AVG PC
WIND SPEED 24H MEAN MS
EVAPORATION MM
RAINFALL 24H MM

The lagged feature order follows the same readable dataset order: X weather variables first and historical rainfall last.

The notebook displays the complete supervised lagged dataset after this step, and each completed script/notebook run saves the same table as lagged_dataset.csv.

Custom datasets use the same column names so the preprocessing and model workflow can run without code changes.

Hyperparameter Search

The project tunes five XGBoost hyperparameters:

colsample_bytree = [0.6, 0.8, 1.0]
learning_rates = [0.05, 0.1, 0.3]
max_depths = [4, 6, 8]
n_estimators = [100, 300, 500]
subsamples = [0.6, 0.8, 1.0]

TimeSeriesSplit is fixed at 10 folds for every run and is not part of the grid search.

Total combinations:

3 x 3 x 3 x 3 x 3 = 243

With 10 TimeSeriesSplit folds, the grid search performs:

243 x 10 = 2430 cross-validation fits

The XGBoost model used for the tuned search is:

XGBRegressor

The default comparison model uses XGBRegressor defaults for the tuned hyperparameters, with the project seed passed internally as random_state for reproducibility, while keeping the same manual-calculation settings from the configuration cell.

Fixed Manual-Calculation Settings

These settings affect the equations used in a manual XGBoost calculation. The objective is fixed by the workflow, while the remaining settings are available in the notebook configuration cell:

Setting Role in manual calculation
Fixed objective: reg:squarederror Defines the gradient and hessian formula for this regression workflow
tree_method Controls how candidate splits are searched (exact, hist, etc.)
base_score Initial prediction before the first tree
reg_lambda L2 regularization term in leaf-weight calculation
reg_alpha L1 regularization term in leaf-weight calculation
gamma Minimum split gain required to create a split
min_child_weight Minimum child-node hessian/weight requirement
max_bin Number of histogram bins when using hist or approx

They are not part of the grid search. They are fixed for a run and applied to every grid-search combination and to the default comparison model. The tuned grid remains:

colsample_bytree, learning_rate, max_depth, n_estimators, subsample

base_score can be configured in three ways:

Setting Meaning
base_score=None Leave the initial prediction to XGBoost's default/automatic handling
base_score=0.5 Use a fixed initial prediction of 0.5 in every TimeSeriesSplit fold
base_score="train_mean" Use mean(y_train_scaled) separately inside each fold

For manual-calculation checks, the notebook reference configuration uses base_score="train_mean" because the calculation starts from the training-target mean. The mean is calculated only from each fold's normalized training target, so the fold test period is not used when setting the initial prediction. Set base_score=None for XGBoost's default/automatic handling.

For a small manual-check setup, reduce the grid to one value per tuned parameter, use CPU, and choose a compact tree:

device = "cpu"
tree_method = "exact"
n_estimators = [1]
max_depths = [1]
learning_rates = [1.0]
subsamples = [1.0]
colsample_bytree = [1.0]
base_score = "train_mean"
reg_lambda = 1.0
reg_alpha = 0.0
gamma = 0.0
min_child_weight = 1.0

Metrics

The main evaluation metric is Normalized MAAPE (%):

MAAPE = mean(arctan2(abs(y_true - y_pred), abs(y_true)))
Normalized MAAPE = MAAPE / (pi / 2)
Normalized MAAPE (%) = Normalized MAAPE * 100

MAAPE is used because rainfall data often contains many zero values, where ordinary MAPE can become unstable. The grid search uses mean CV Normalized MAAPE (%) for model selection.

arctan2 is used instead of arctan(abs(error) / abs(actual)) because it avoids manual division by the actual rainfall value. This keeps the metric defined when the actual rainfall is 0. When both actual and predicted rainfall are 0, the MAAPE contribution is 0; when the actual rainfall is 0 but the prediction is not 0, the contribution becomes the maximum arctangent penalty.

The normalization step converts the original MAAPE angle from the range 0 to pi/2 into a more readable 0% to 100% scale. The original MAAPE angle is also reported as a supporting metric, while MAE and RMSE are reported in millimeters. Metric tables are displayed with 4 decimal places in the notebook.

Result Interpretation

  • Lower Normalized MAAPE (%), MAAPE, MAE, and RMSE values indicate better predictive performance.
  • The best tuned model is selected using mean CV Normalized MAAPE (%), not MAE or RMSE. Plain MAAPE is displayed only as an additional metric.
  • MAE and RMSE are reported in millimeters, so they are easier to interpret in the original rainfall unit.
  • A tuned model can improve Normalized MAAPE (%) while not improving every supporting metric. This can happen because each metric penalizes errors differently.
  • The final prediction CSV contains out-of-fold predictions. Each row is predicted by a model that was trained only on data before that fold's test period.
  • Results produced with sample_weather_data.csv are intended for reproducible public demonstration. They are not interpreted as the performance of private raw data.

Post-Processing

Rainfall is a non-negative variable, so negative model predictions are clipped to 0 after predictions are denormalized back to millimeters. This clipping is applied before evaluation/scoring, CSV export, and plotting.

Limitations

  • The dataset represents one observation area, so the result is not assumed to generalize directly to other regions without retraining or further testing.
  • Rainfall data contains many zero-rainfall days and sudden high-rainfall events. The model may follow the general pattern better than sharp extreme peaks.
  • The default run uses a 1-day lag window. Other lag lengths may produce different results and can be tested from the notebook configuration cell.
  • The experiment uses non-nested TimeSeriesSplit cross-validation. The reported CV score is useful for comparing hyperparameter combinations, but it may still be slightly optimistic because the same CV process is used for tuning and reporting.
  • The earliest initial training segment does not have out-of-fold predictions because TimeSeriesSplit preserves chronological order.
  • The model only uses the variables available in the CSV file. Additional predictors such as radar, satellite, numerical weather prediction, or broader climate indices may improve performance if available.
  • Exact results can vary slightly across package versions and runtime environments, although the code sets a random seed for reproducibility.
  • Linear interpolation is applied as an offline data-cleaning rule and can use observations on both sides of a gap. A real-time operational pipeline should replace it with a causal imputation rule that only uses information available before the forecast timestamp.

Installation

Python 3.12 is recommended. The code requires Python 3.10 or newer.

pip install -r requirements.txt

The same requirements file is used for CPU and optional NVIDIA/CUDA GPU runs. CPU users can keep device="cpu" or device="auto" in the notebook. GPU users use device="gpu" with tree_method="hist".

Usage

Run the notebook:

XGBoost_Rainfall_Prediction.ipynb

Check preprocessing without training:

python train_xgboost_rainfall.py --prepare-only

Run full training:

python train_xgboost_rainfall.py

Force CPU:

python train_xgboost_rainfall.py --device cpu

Request GPU:

python train_xgboost_rainfall.py --device gpu

Show the final plot after training:

python train_xgboost_rainfall.py --show-plot

Run a smaller grid search:

python train_xgboost_rainfall.py --colsample-bytree 0.8 --learning-rates 0.1 --max-depths 4 --n-estimators 100 --subsamples 0.8

Reproducibility and Runtime

  • The notebook and script default to the included synthetic sample dataset, so the public workflow can run without private data.
  • A random seed is set for reproducibility, but exact results can still vary slightly across package versions and runtime environments.
  • Full XGBoost search evaluates 243 hyperparameter combinations across 10 TimeSeriesSplit folds. Train-fold metrics and detailed out-of-fold predictions are calculated only for the selected combination, avoiding unnecessary train inference for every candidate. The selected and untuned comparison setups are then evaluated once on the same folds for their detailed outputs.
  • Use --prepare-only to verify data loading and preprocessing without model training.
  • Use the smaller grid-search command above for a faster functional check.

Outputs

Each completed run creates an output folder under outputs/. By default, only the latest completed run folder is kept.

Main output files:

File Description
lagged_dataset.csv Complete supervised lagged dataset used for modeling
hyperparameter_results.csv All grid-search combinations sorted by mean_cv_normalized_maape_percent
best_cv_fold_metrics.csv Fold-by-fold metrics for the best grid-search model
best_cv_predictions.csv Out-of-fold actual vs predicted rainfall for the best model
best_cv_prediction_plot.png Actual vs predicted out-of-fold plot for the best model
default_cv_fold_metrics.csv Fold-by-fold metrics from the default model
default_cv_predictions.csv Out-of-fold actual vs predicted rainfall from the default model
default_cv_prediction_plot.png Default model out-of-fold plot
best_vs_default_comparison.csv Best grid-search model vs default model
run_metadata.json Configuration, compute runtime, feature columns, split information, and preprocessing details

The notebook also displays a compact final summary:

Best model Normalized MAAPE (%):
Best model MAAPE:
Best model MAE:
Best model RMSE:
Default model vs tuned model:

Generated outputs, model artifacts, Python caches, and notebook checkpoints are ignored by Git.

License

This project is released under the MIT License. The license applies to the project code and the included synthetic sample dataset. Private raw datasets are not redistributed in this repository and are not covered as public datasets by this license.

About

Multivariate daily rainfall forecasting with XGBoost regression, lag features, TimeSeriesSplit cross-validation, and zero-safe Normalized MAAPE.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages