This is a technical test to assess the engineering and solutioning competencies for a senior data engineer role in HDB.
Requires Python 3.12+.
- Create and activate a virtual environment:
python -m venv .venv
source .venv/bin/activate- Install dependencies:
pip install -r requirements.txt -r requirements-dev.txt- Register the venv as a Jupyter kernel:
python -m ipykernel install --user --name hdb-etl- Select the
hdb-etlkernel in Jupyter before running the code in submission.ipynb.
Steps to execute the ETL pipeline are outlined in the submission.ipynb Jupyter notebook.
The following is a proposed architecture that builds on top of the AWS cloud stack. The full design which includes assumptions, data ingestion and exploitation flows, and security/scalability/performance considerations is documented in architecture.md.
