Skip to content

parallelworks/activate-cluster-demos

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

7 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ACTIVATE GPU Cluster: Quick Start and Demos

Runnable examples and a command reference for the NVIDIA H100 GPU cluster operated by Parallel Works on dedicated CoreWeave infrastructure and managed through the ACTIVATE control plane. Everything here has been run on the live cluster. Questions go to support@parallelworks.com.

Clone this repo onto the cluster and run the examples from your home directory:

git clone https://github.com/parallelworks/activate-cluster-demos.git
cd activate-cluster-demos
Folder What it shows
01-pytorch Modules and a single- to multi-GPU PyTorch job
02-containers Containers through Slurm with pyxis and enroot
03-jupyter A Jupyter notebook on a GPU node over an SSH tunnel
04-ray A Ray cluster on Slurm, with jobs submitted from your laptop
05-k8s Kubernetes (CKS) namespace access and quotas
06-qos How GPUs are shared, mylimits / myquota, and a runnable example-job.sh

Getting on the cluster

Your account is provisioned in ACTIVATE at activate.parallel.works; set your password from the email invitation. From a terminal, install the pw CLI and connect with pw ssh nodescluster. VS Code and remote desktop sessions are also available through ACTIVATE.

How the scheduler shares GPUs

The cluster runs Slurm. There is a single GPU partition, h100, and it is the default, so you do not have to name it; including -p h100 is fine and harmless.

Every group has a group QOS, called your Project QOS, and it is your default. It caps the number of GPUs your group can run at once at the number your team was allocated; you never type it. On top of that, --qos=general bursts into a shared pool up to a per-group limit, and --qos=spot is uncapped and preemptible: spot jobs run on idle GPUs and can be reclaimed with a 10-minute grace window, then requeue, so checkpoint your work. Project and general jobs are never preempted. To see your own QOS, cap, and current usage, run mylimits (or myquota, the same command), which is installed on the cluster. 06-qos/README.md documents the underlying sacctmgr commands.

Run a PyTorch job

module load pytorch
srun -p h100 --gpus=1 python 01-pytorch/gputest.py

Returns cuda_available: True and NVIDIA H100 80GB HBM3. 01-pytorch/train.sh runs the 8-GPU torchrun form against train.py (a few steps on synthetic data, so it runs as-is; output lands in train-<jobid>.out); multi-node adds --nodes=2 --gpus-per-node=8, with NCCL over InfiniBand already tuned. The same pattern runs NAMD, OpenMM, and JAX: load the module, then srun or sbatch.

Ready-made job scripts for the common cases live in /software/templates/ on the cluster.

Containers with pyxis and enroot

bash 02-containers/pyxis_hello.sh      # two verified one-liners
bash 02-containers/enroot_build.sh     # imports, builds, and runs on a GPU node

Use glibc images such as ubuntu:22.04 or the NGC catalog, not ubuntu:latest or alpine. enroot import runs on a compute node, and create/start must run on the same node. Large images pull once and then cache on the node.

Jupyter on a GPU node

sbatch 03-jupyter/jupyter.sh
squeue --me                            # note the node
grep -o "http://.*token=.*" jupyter-<jobid>.out | head -1

On your laptop, open a tunnel and browse to the notebook:

pw ssh nodescluster -L 8888:<gpu-node>:8888
# open http://localhost:8888 and paste the token

Ray, submitting from your laptop

sbatch 04-ray/raycluster.sh
squeue --me                            # note the head node

On your laptop (with pip install ray and the pw CLI):

pw ssh nodescluster -L 8265:<head-node>:8265
export RAY_ADDRESS=http://localhost:8265
ray job submit --working-dir 04-ray -- python train_torch.py
# -> ('NVIDIA H100 80GB HBM3', True)  SUCCEEDED

Where your data lives

Your home /home/<username> (small quota) holds code and environments; your group's /project/<group> (large, shared, on every node) holds datasets and results; /software is the Lmod module tree. Each GPU server also has fast node-local NVMe scratch that is ephemeral, so copy results back to your project directory.

Troubleshooting

module: command not found, or python not found in a job. Modules now work in batch scripts as well as interactive sessions: the cluster sets MODULEPATH and defines module for non-login shells, so a plain #!/bin/bash script with module load pytorch works, and the scripts here and in /software/templates/ rely on that. Do not add -l to the shebang; on this cluster a login shell prints a spurious FATAL: Missing required source file. line into your job output.

If you do hit a missing module, it is almost always one of two things. For a non-interactive command over SSH, wrap it in a login shell: pw ssh nodescluster "bash -lc 'module load pytorch; python train.py'". And watch out for pipes: module load pytorch | tee log runs the load in a subshell, so nothing is loaded in your script. Load first, then run.

srun says execve(): python: No such file or directory. The module was not loaded in the environment srun inherited. Run module load pytorch in the same shell first, then srun, and confirm with command -v python; it should be under /software.

A job sits in PENDING with reason QOSGrpGRES. That is your GPU cap, not an error. The job starts as your running jobs finish. See 06-qos for how to view your limits.

Module ERROR: Magic cookie '#%Module' missing. Cosmetic Lmod noise from a stray file in the module tree; it does not affect your job. Report it to support and we will clear it.

Getting help

support@parallelworks.com is staffed 24/7. There is a standing weekly office-hours session, and we run additional training on request. User Guide and CLI Reference: docs.parallel.works.

About

Quick start and runnable demos for the Parallel Works ACTIVATE GPU cluster (Slurm, containers, Jupyter, Ray, Kubernetes) on CoreWeave H100s.

Resources

Stars

Watchers

Forks

Releases

Packages

Used by

Contributors

Languages