I turn messy, real-world data into clear, decision-ready analysis — in Python, with reproducible pipelines.
My path into data analytics
1 · Foundations — Python, pandas, EDA, data cleaning
- Data cleaning & validation — real dataset with errors and ~50% missing values; every cleaning decision documented and validated before/after.
- Multi-table integration — MovieLens —
concat/mergeacross 4 tables, each join validated for row/null integrity.
2 · Applied analysis & visualization — EDA and dashboards for real decisions
- Car sales — EDA & dashboard — seasonality and market share, 6-panel dashboard.
- Bird migration — comparative EDA and time series, interactive dashboard in Tableau Public.
3 · Bioinformatics & scientific data — real sequencing data, tested code, publication-level rigor
- Microbial diversity — ENIGMA study — 67 groundwater wells, 16S rRNA data, self-audited against the STREAMS reporting standard, tested Python package.
- ITS amplicon pipeline — DADA2/cutadapt pipeline reproducing the nf-core/ampliseq workflow, automated tests, CI.
Next: M.S. in Data Analytics (starting 2026).
Also building Verdika, a Decision Intelligence platform.