- Run Models (CoLab) - For features and targets merged on their location columns.
- Models Overview
In Run Models, the "features" dataset is merged with a 2-column "targets" dataset on-the-fly using either .csv files or Pandas to avoid storing merged .csv files. The location column joins features and targets.
We also support feature files that already contain a target column, like the eye blink data Random Bits Forest (RBF).
Location column data types:
World Region (TBD), Country (2-char), State (2-char), County Fips (5-digits for state and county), Zip (5 char, 6 in China), or Brain Voxel (2 char)
The target does not need to a location ID. It can be an ID that clusters multiple location, as is the case for Brain Voxels that fire together when eye blinks occur. View eye blink data .csv file - each column is a voxel (location in brain).
Our features-targets merge supports any data with a location column containing multiple locations.
Our default data will always use County Fips so features and targets align.
Industries (Features and Targets) - County Fips Industries Input Data
Bees (Target) - County Fips Random Forest (Bees)
Trees (Target) - County Fips Tree Targets
Blinks - Rows are clusters of brain voxels - hence multiple locations have one target column
Random Bits Forest (Blinks)
You can add paths to external data by editing a copy of the parameters.yaml file.
The term "features" is more prevalent in machine learning and data science. "factors" has a stronger association with statistics and social sciences. The term factors is used for impact attributes like emissions.
The following python command loads parameters.yaml to run Run-Models-bkup.ipynb locally without opening a notebook (since edits are saved in the Google colab rather than the local backup).
Parameters are loaded into Run-Models-bkup.ipynb from the parameters.yaml file. TO DO: Debug errors. Edit Google colab and save to bkup file.
python run_models_cli.py parameters/parameters.yaml
This one doesn't run the Run-Models-bkup.ipynb file yet:
python run_models.py parameters/parameters.yaml
Example of parameters.yaml format:
folder: naics6-bees-counties
features: industries
startyear: 2017
endyear: 2021
path: https://raw.githubusercontent.com/ModelEarth/community-timelines/main/training/naics{naics}/US/counties/{year}/US-{state}-training-naics{naics}-counties-{year}.csv
targets: bees
path: https://github.com/ModelEarth/bee-data/raw/main/targets/bees-targets-top-20-percent.csv
models: lr, svc, rfc, rbf, xgboost
Each target dataset will contain 2 columns.
- The location column with one of the following column names:
Country (2-char), State (2-char), Fips (5-digits for state and county), Zip (5 char, 6 in China), or Voxel (2 char) - The "Target" column containing 1 or 0
Setting the models parameter to "all" would be the equivalent to "lr,rfc,rbf,svm,mlp,xgboost"
Default features and targets datasets reside in the "input/[data]/features" and "input/[data]/targets" folders for each data source.
The simplest form of the parameters.yaml would be:
features: industries
targets: bees
That's the equivalent to:
features: industries
path: https://github.com/ModelEarth/realitystream/raw/main/input/industries/features/industries-features.csv
targets: bees
path: https://github.com/ModelEarth/bee-data/raw/main/targets/bees-targets-top-20-percent.csv
models: rbf
The features.path and targets.path will have several shorthand versions and a full version from GitHub:
short - bees
medium - /bee-data/targets
long - /bee-data/targets/bees-targets-top-20-percent.csv
full - https://github.com/ModelEarth/bee-data/raw/main/targets/bees-targets-top-20-percent.csv
Path processing rules: If there's no slash / in a path parameter, start from the root of the RealityStream repo. If the file extension is omitted from a path, append .csv. For a target value of "bees" build the path "/bee-data/targets/bees-targets-top-20-percent.csv" Replace a space with -targets- in the path. So for a target value of "bees increase2024" build the path "/bee-data/targets/bees-targets-increase2024.csv"