Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

雀巢咖啡消费者大数据系统(毕业设计 — 系统部分)

本仓库为系统架构与实现部分,仅使用合成数据,不包含真实业务数据。

文档

  • 系统架构设计说明(推荐):系统整体架构详细说明,含逻辑分层、数据湖、计算栈、四大模型、Serve 层、营销对接、安全与部署。
  • 系统设计说明:总体架构、数据湖、双栈计算、四大模型、营销对接与 A/B、CI 与安全。
  • 双栈计算平台说明:实时栈与离线栈组件与数据流。
  • 调度与运行说明:run_all / run_daily、crontab、Airflow、Docker、营销对接与 A/B 入口。
  • 数据获取方案:各端数据获取方式、与本系统对接、CSV 导入 Raw 层、公开数据集与合规。

项目结构

coffee/
├── config/             # 配置(数据湖路径、Spark、合成数据规模)
├── src/                # 公共模块(表结构、Spark 工具)
├── pipeline/            # 数据接入与 ETL、四大模型、营销对接
│   ├── ingest_synthetic.py   # 合成数据写入 Raw(含 is_promotion)
│   ├── etl_raw_to_cleaned.py # Raw → Cleaned
│   ├── etl_cleaned_to_model.py# Cleaned → Model 用户特征
│   ├── model_user_profile.py # 用户画像 → Serve
│   ├── model_user_cluster.py # 细分聚类(KMeans)→ Serve
│   ├── model_price_elasticity.py # 价格弹性 → Serve
│   ├── model_basket.py        # 购物篮关联规则 → Serve
│   ├── serve_audience_export.py # 人群包导出(营销对接)
│   ├── ab_experiment_config.py  # A/B 实验配置
│   ├── ab_significance.py       # A/B 显著性检验
│   └── run_all.py              # 一键跑通整条流水线
├── scripts/
│   └── run_daily.py    # 按日调度(默认昨日)
├── data/               # 数据湖根目录(运行后生成)
├── tests/
├── docs/
├── Dockerfile
├── requirements.txt
└── README.md

环境要求

  • Python 3.10+
  • JDK 17+(PySpark 需要)
  • 建议在项目根目录下执行所有命令

安装与运行

1. 虚拟环境与依赖(推荐)

cd /path/to/coffee
python3 -m venv .venv
source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install -r requirements.txt

2. 一键跑通离线流水线(合成数据)

export PYTHONPATH=.
python -m pipeline.run_all

可选指定日期(默认当天):

python -m pipeline.run_all 2024-01-15

流水线顺序:Raw(合成)→ Cleaned → Model → Serve(用户画像 + 细分聚类 + 价格弹性 + 购物篮)。结果落在 data/serve/user_profile/、user_cluster/、price_elasticity/、basket_rules/ 等。

3. 运行单元测试

PYTHONPATH=. pytest tests/ -v

4. Docker 运行(可选)

docker build -t coffee-pipeline .
docker run --rm coffee-pipeline

镜像内使用合成数据执行整条流水线,数据写入容器内 data/,可用于演示与交付。

CI

推送至 main / master 后,GitHub Actions 执行:安装依赖 → 运行 pytest tests/。无真实数据,仅代码与逻辑校验。

系统要点摘要

  1. 数据湖:Raw → Cleaned → Model → Serve 四层;交易端 + 行为端合成数据(含 unit_price、is_promotion)。
  2. 四大模型:用户画像、细分聚类(KMeans)、价格弹性、购物篮关联规则;结果均落 Serve 层。
  3. 营销对接与 A/B:人群包导出(按 segment)、实验配置表、显著性检验(t 检验)写 ab_result。
  4. 调度:run_all 一键跑通;scripts/run_daily.py 按日跑昨日,可接 crontab / Airflow。
  5. 安全:全流程脱敏、无 PII;白皮书/镜像/手册不出现真实数据。
  6. 可扩展:实时栈(Kafka+Flink)、需求预测模型、与雀巢营销系统 API 对接。

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages