Skip to content

Repository files navigation

Introduction

This is a technical test to assess the engineering and solutioning competencies for a senior data engineer role in HDB.

Setup

Requires Python 3.12+.

  1. Create and activate a virtual environment:
python -m venv .venv
source .venv/bin/activate
  1. Install dependencies:
pip install -r requirements.txt -r requirements-dev.txt
  1. Register the venv as a Jupyter kernel:
python -m ipykernel install --user --name hdb-etl
  1. Select the hdb-etl kernel in Jupyter before running the code in submission.ipynb.

Part 1: Developing Data Pipelines

Steps to execute the ETL pipeline are outlined in the submission.ipynb Jupyter notebook.

Part 2: Architecting Data Ingestion & Data Exploitation Solution Patterns

The following is a proposed architecture that builds on top of the AWS cloud stack. The full design which includes assumptions, data ingestion and exploitation flows, and security/scalability/performance considerations is documented in architecture.md.

Solution architecture diagram

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages