本仓库为系统架构与实现部分,仅使用合成数据,不包含真实业务数据。
- 系统架构设计说明(推荐):系统整体架构详细说明,含逻辑分层、数据湖、计算栈、四大模型、Serve 层、营销对接、安全与部署。
- 系统设计说明:总体架构、数据湖、双栈计算、四大模型、营销对接与 A/B、CI 与安全。
- 双栈计算平台说明:实时栈与离线栈组件与数据流。
- 调度与运行说明:run_all / run_daily、crontab、Airflow、Docker、营销对接与 A/B 入口。
- 数据获取方案:各端数据获取方式、与本系统对接、CSV 导入 Raw 层、公开数据集与合规。
coffee/
├── config/ # 配置(数据湖路径、Spark、合成数据规模)
├── src/ # 公共模块(表结构、Spark 工具)
├── pipeline/ # 数据接入与 ETL、四大模型、营销对接
│ ├── ingest_synthetic.py # 合成数据写入 Raw(含 is_promotion)
│ ├── etl_raw_to_cleaned.py # Raw → Cleaned
│ ├── etl_cleaned_to_model.py# Cleaned → Model 用户特征
│ ├── model_user_profile.py # 用户画像 → Serve
│ ├── model_user_cluster.py # 细分聚类(KMeans)→ Serve
│ ├── model_price_elasticity.py # 价格弹性 → Serve
│ ├── model_basket.py # 购物篮关联规则 → Serve
│ ├── serve_audience_export.py # 人群包导出(营销对接)
│ ├── ab_experiment_config.py # A/B 实验配置
│ ├── ab_significance.py # A/B 显著性检验
│ └── run_all.py # 一键跑通整条流水线
├── scripts/
│ └── run_daily.py # 按日调度(默认昨日)
├── data/ # 数据湖根目录(运行后生成)
├── tests/
├── docs/
├── Dockerfile
├── requirements.txt
└── README.md
- Python 3.10+
- JDK 17+(PySpark 需要)
- 建议在项目根目录下执行所有命令
cd /path/to/coffee
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txtexport PYTHONPATH=.
python -m pipeline.run_all可选指定日期(默认当天):
python -m pipeline.run_all 2024-01-15流水线顺序:Raw(合成)→ Cleaned → Model → Serve(用户画像 + 细分聚类 + 价格弹性 + 购物篮)。结果落在 data/serve/user_profile/、user_cluster/、price_elasticity/、basket_rules/ 等。
PYTHONPATH=. pytest tests/ -vdocker build -t coffee-pipeline .
docker run --rm coffee-pipeline镜像内使用合成数据执行整条流水线,数据写入容器内 data/,可用于演示与交付。
推送至 main / master 后,GitHub Actions 执行:安装依赖 → 运行 pytest tests/。无真实数据,仅代码与逻辑校验。
- 数据湖:Raw → Cleaned → Model → Serve 四层;交易端 + 行为端合成数据(含 unit_price、is_promotion)。
- 四大模型:用户画像、细分聚类(KMeans)、价格弹性、购物篮关联规则;结果均落 Serve 层。
- 营销对接与 A/B:人群包导出(按 segment)、实验配置表、显著性检验(t 检验)写 ab_result。
- 调度:
run_all一键跑通;scripts/run_daily.py按日跑昨日,可接 crontab / Airflow。 - 安全:全流程脱敏、无 PII;白皮书/镜像/手册不出现真实数据。
- 可扩展:实时栈(Kafka+Flink)、需求预测模型、与雀巢营销系统 API 对接。