Build a small virtual HPC cluster on local macOS with Docker, to learn Slurm architecture, configuration, administration, and troubleshooting.
Current status: Phase 1–3 complete — MPI cross-node, OpenMP, Lmod module system all verified.
┌──────────────────────────────────────────────────────────────┐
│ Docker Network │
│ 172.20.0.0/16 (cluster-net) │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ slurmctl │ │ node01 │ │ node02 │ │
│ │ (controller) │ │ (compute) │ │ (compute) │ │
│ │ 172.20.0.10 │ │ 172.20.0.11 │ │ 172.20.0.12 │ │
│ │ slurmctld │ │ slurmd │ │ slurmd │ │
│ │ mariadb │ │ │ │ │ │
│ │ munge │ │ munge │ │ munge │ │
│ │ nfs-server │ │ nfs-client │ │ nfs-client │ │
│ │ sshd │ │ sshd │ │ sshd │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
└──────────────────────────────────────────────────────────────┘
- slurmctl — Controller: slurmctld + slurmdbd + mariadb + NFS server
- node01/node02 — Compute: slurmd + NFS client (/home /scratch) + SSH
cluster-lab/
├── README.md # This file
├── LICENSE # MIT License
├── .gitignore # Docker / IDE / OS ignore rules
├── docker/
│ ├── Dockerfile # Unified image (controller + compute)
│ ├── docker-compose.yml # Multi-container orchestration
│ └── entrypoint.sh # Container entrypoint by role
├── config/
│ ├── slurm/
│ │ ├── slurm.conf # Slurm main configuration
│ │ ├── slurmdbd.conf # SlurmDBD configuration
│ │ └── cgroup.conf # cgroup resource isolation
│ ├── nfs/
│ │ └── exports # NFS share definitions
│ └── modulefiles/ # Lmod modulefiles
│ ├── gcc/
│ │ └── 13.3.0.lua # GCC module definition
│ └── openmpi/
│ └── 4.1.6.lua # OpenMPI module definition
├── scripts/
│ └── setup.sh # One-click deploy (check → build → up → verify)
├── tests/
│ ├── hello_mpi.c # MPI parallel test (cross-node capable)
│ └── hello_openmp.c # OpenMP shared-memory test
├── 日记/ # Project diaries (Chinese)
│ ├── 日记-项目架构与设计笔记.md # Architecture & design decisions
│ ├── 日记-部署阶段实战笔记.md # 11 deployment pitfalls & fixes
│ └── 日记-跨节点MPI与Lmod配置.md # MPI cross-node & module system
└── journal/ # (planned) Experiment logs
# 1. Build the image
docker compose -f docker/docker-compose.yml build
# 2. Start the cluster
docker compose -f docker/docker-compose.yml up -d
# 3. Check node status
docker exec slurmctl sinfo
# 4. Submit a test job
docker exec slurmctl srun -N 2 hostname
# 5. Run MPI cross-node
docker exec slurmctl srun -N 2 -n 2 --mpi=pmix /home/hello_mpi| Capability | Status | Verification |
|---|---|---|
| Cluster startup (3 nodes) | ✅ | docker ps + sinfo |
| Slurm job scheduling | ✅ | srun hostname |
| NFS shared filesystem | ✅ | Cross-node /home access |
| OpenMP parallel computing | ✅ | ./hello_openmp |
| MPI single-node parallel | ✅ | srun -n 2 ./hello_mpi |
| MPI cross-node (srun) | ✅ | srun -N 2 -n 2 --mpi=pmix |
| MPI cross-node (SSH) | ✅ | mpirun --host node01,node02 -np 2 |
| Lmod module environment | ✅ | module load gcc/13.3.0 openmpi/4.1.6 |
| Git version control | ✅ | GitHub: xczics/cluster-lab |
| Phase | Status | What |
|---|---|---|
| Phase 1: Deployment | ✅ | 11 issues resolved — Docker, Slurm, NFS, Munge, cgroup, network |
| Phase 2: Compilers & MPI | ✅ | OpenMP 4.5 + OpenMPI 4.1.6 installed and verified |
| Phase 3: Module System | ✅ | Lmod — module load gcc/13.3.0 + openmpi/4.1.6 |
| Phase 4: Cross-node MPI | ✅ | A+B dual approach — PMIx (srun) + SSH (mpirun) |
mpirun --host node01,node02 -np 2 /home/hello_mpiIndependent of Slurm, uses SSH keys for node-to-node authentication.
srun -N 2 -n 2 --mpi=pmix /home/hello_mpiFully integrated with Slurm scheduler — resource isolation, job queuing, priority.
Hello from MPI process 1 of 2 on node02
Hello from MPI process 0 of 2 on node01
| Component | Detail |
|---|---|
| Host | macOS (Apple Silicon / Intel) |
| Container | Docker — Ubuntu 24.04 LTS |
| Scheduler | Slurm 24.11.3 (apt) + PMIx integration |
| MPI | OpenMPI 4.1.6 — Built with --with-pmix |
| OpenMP | GCC 13.3.0 built-in |
| Module System | Lmod 8 (Lua 5.4) |
| Auth | Munge (slurmctld → slurmd) |
| Storage | NFS v3 (nolock) — /home + /scratch |
| Network | Docker custom bridge — 172.20.0.0/16 |
- Controller node (slurmctl): slurmctld + slurmdbd + mariadb + NFS server + sshd
- Compute nodes (node01/node02): slurmd + NFS client + sshd
- Container image: Single Dockerfile, role selected by
SLURM_ROLEenv var - cgroup: v1 with
linuxproctracking (Docker Desktop compatibility) - Accounting: Disabled (
AccountingStorageType=none) — Munge errors bypassed - All nodes: Cross-mounted
/homevia NFS, shared SSH authorized_keys
| # | Feature | Files | Status |
|---|---|---|---|
| 1 | sbatch job templates | scripts/jobs/hello_mpi.sbatch, hello_openmp.sbatch, README.md |
✅ sbatch-ready |
| 2 | Cluster validation script | scripts/test-cluster.sh (11 checks) |
✅ bash /home/scripts/test-cluster.sh |
| 3 | Multi-user environment | scripts/setup_users.sh (alice/bob/charlie) |
✅ bash /home/scripts/setup_users.sh |
| 4 | Slurm QoS | config/slurm/qos_partitions.conf, scripts/setup_qos.sh (high/normal/low) |
✅ bash /home/scripts/setup_qos.sh |
| 5 | GRES GPU simulation | config/slurm/gres.conf + tests/check_gpu.c |
✅ 4 fake GPUs per node |
| 6 | Job array & dependency templates | scripts/jobs/job_array.sbatch, job_deps.sbatch |
✅ sbatch --array=1-10 / sbatch --dependency=afterok:$JOBID |
| 7 | Container multi-stage build | docker/Dockerfile.multistage |
✅ ~1.5GB → ~600-800MB |
| 8 | Ansible automation | ansible/site.yml + inventory.yml |
✅ ansible-playbook -i ansible/inventory.yml ansible/site.yml |
| 9 | Cluster health monitoring | scripts/monitoring/cluster_health.sh |
✅ bash scripts/monitoring/cluster_health.sh |
- Prometheus + Grafana dashboard
- CI/CD with GitHub Actions
- Docker Hub image publishing
- Production Munge authentication
MIT