A one-screen terminal dashboard for small GPU clusters. CPU, RAM, per-GPU utilisation for every node, plus the Slurm queue โ with gradient bars and sparkline history, in a single 24-line frame.
Built during HiPAC 2026 because watch nvidia-smi
on each node in a separate tmux pane got old fast.
A 2-node / 16รH200 cluster going from idle to pegged and back. The load here
is simulated โ same render path, synthetic telemetry โ because the cluster was
powered down when this was recorded. Everything you see is what the real thing
draws: the scope filling, both node zones going red, heat plumes at 70 ยฐC,
FULL LOAD lighting up, and the progress bars creeping toward each job's
time limit.
Same thing as text
โค SLURMTOP // hipac-team3 โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ 11:56:32 โฅ
โโโโ
โโโโโโโโ
โโโ CORE 16/16 online โ FULL LOAD
99.2% UTIL PWR 10.49 kW MEM 1917/2246 GiB THRM 70ยฐC
ยท โโโโโโโโโโโโ GRID โโโโโโโโ โ โโโโโโโโ
โญโค LOAD // last 63 samples โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ ยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทโ
โ
โ
โโโโโโโโโโโโโโโโโโโโโโโโยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยท โ
โ ยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยท โ
โ ยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยท โ
โ ยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทโ
โ
โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยท โ
โ ยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทโโโโ
โ
โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยท โ
โ ยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยทยท โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโค NODE n1 โ LOCAL โโโโโโโโโโโโโโโโโโโโโ 8/8 busy ยท 958G ยท 5.2 kW ยท 70ยฐC โฎโญโค NODE n2 โ SSH โโโโโโโโโโโโโโโโโโโโโโโ 8/8 busy ยท 958G ยท 5.2 kW ยท 62ยฐC โฎ
โ CPU โโโโโโโโโโโโโโโโ 8.0% ยท โโโโโโโโ load 18 8 7 โโ CPU โโโโโโโโโโโโโโโโ 8.0% ยท โโโโโโโโ load 18 8 7 โ
โ MEM โโโโโโโโโโโโโโโโ 47.3% 953/2016 GiB โโ MEM โโโโโโโโโโโโโโโโ 47.3% 953/2016 GiB โ
โ G0 โโโโโโโโโโโโโโ 100%~ โโโโโโโโ 119.8G 68ยฐ 660W โโ G0 โโโโโโโโโโโโโโ 100%โ โโโโโโโโ 119.8G 60ยฐ 660W โ
โ G1 โโโโโโโโโโโโโโ 98%โ โโโโโโโโ 119.8G 69ยฐ 648W โโ G1 โโโโโโโโโโโโโโ 98%โ โโโโโโโโ 119.8G 61ยฐ 648W โ
โ G2 โโโโโโโโโโโโโโ 100%โ โโโโโโโโ 119.8G 70ยฐ 660W โโ G2 โโโโโโโโโโโโโโ 100%~ โโโโโโโโ 119.8G 62ยฐ 660W โ
โ G3 โโโโโโโโโโโโโโ 100%~ โโโโโโโโ 119.8G 68ยฐ 660W โโ G3 โโโโโโโโโโโโโโ 100%โ โโโโโโโโ 119.8G 60ยฐ 660W โ
โ G4 โโโโโโโโโโโโโโ 98%โ โโโโโโโโ 119.8G 69ยฐ 648W โโ G4 โโโโโโโโโโโโโโ 98%โ โโโโโโโโ 119.8G 61ยฐ 648W โ
โ G5 โโโโโโโโโโโโโโ 100%โ โโโโโโโโ 119.8G 70ยฐ 660W โโ G5 โโโโโโโโโโโโโโ 100%~ โโโโโโโโ 119.8G 62ยฐ 660W โ
โ G6 โโโโโโโโโโโโโโ 100%~ โโโโโโโโ 119.8G 68ยฐ 660W โโ G6 โโโโโโโโโโโโโโ 100%โ โโโโโโโโ 119.8G 60ยฐ 660W โ
โ G7 โโโโโโโโโโโโโโ 98%โ โโโโโโโโ 119.8G 69ยฐ 648W โโ G7 โโโโโโโโโโโโโโ 98%โ โโโโโโโโ 119.8G 61ยฐ 648W โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏโฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโค Slurm queue โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ โค SCAN โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ JOB NAME STATE ELAPSED PROG LEFT N CPU GPU WHERE โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ 871 hawks-mxp16 โ run 12m32s โโโโโโโโ 47m28s 2 448 8 n[1-2] โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ 872 hawks-qe-pw โ run 9m24s โโโโโโโโ 20m36s 1 224 8 n2 โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ 873 hawks-eagle โ pend 0s โโโโโโโโ 2h00m 1 32 4 (Resources) โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ 874 hawks-cfd โ pend 0s โโโโโโโโ 3h00m 1 14 - (Dependency) โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ n1 allocated CPU โโโโโโโโโโโโ 224/224 idle 0 RAM free 742 GiB โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ n2 allocated CPU โโโโโโโโโโโโ 224/224 idle 0 RAM free 742 GiB โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ github.com/Sean-Hawks/slurmtop
Ctrl-C quit ยท --proc processes ยท --stack vertical ยท -n <sec> interval team-03/Hawks ยท v1.0.0
The whole panel is tinted by node load: n1 pegged and hot, n2 middling, idle nodes stay dark.
nvtop is great but shows one machine. squeue tells you what is queued but
not whether the GPUs are actually doing anything. On a handful of nodes you
usually want both at once, on one screen, without installing an agent, a time
series database and a web UI.
Single file, no dependencies beyond the standard library.
curl -fsSL https://raw.githubusercontent.com/Sean-Hawks/slurmtop/main/install.sh | bashThe installer picks /usr/local/bin when it can write there (using sudo if
passwordless sudo is available) and falls back to ~/.local/bin, telling you
how to fix your PATH if needed. Pass a directory to choose yourself:
curl -fsSL .../install.sh | bash -s -- ~/binOr just drop the single file in place โ that is all the installer does:
curl -fsSLo /usr/local/bin/slurmtop \
https://raw.githubusercontent.com/Sean-Hawks/slurmtop/main/slurmtop
chmod +x /usr/local/bin/slurmtopOnly the machine you run it from needs the script โ it reads the other nodes over SSH. Installing it on every node is optional but handy.
Requirements:
- Python 3.8+ on the machine you run it from
- Linux or macOS on each node
- passwordless SSH from that machine to every remote node
Everything else is optional and degrades cleanly:
| Missing | What happens |
|---|---|
nvidia-smi / no GPU |
node panels show CPU and RAM only, and the header switches its main gauge to CPU |
| Slurm | pass --nodes; the queue panel just says there are no jobs |
| a second machine | slurmtop --nodes localhost watches the box you are on, no SSH involved |
CPU and memory are read from /proc on Linux and from sysctl / vm_stat /
top on macOS, so a laptop works as a node like anything else. Only NVIDIA
GPUs are read; AMD and Intel are not supported yet โ the coupling is one
nvidia-smi call in REMOTE, so a patch adding rocm-smi would be small.
slurmtop # refresh every 2s, nodes discovered from Slurm
slurmtop -n 5 # refresh every 5s
slurmtop --once # print one frame and exit (good for chat/logs)
slurmtop --nodes a,b,c # explicit node list, skip Slurm discovery
slurmtop --nodes localhost # single machine, no SSH, no Slurm needed
slurmtop --proc # also list the processes on each GPU
slurmtop --stack # force vertical layout
slurmtop --fit # squeeze into one screen instead of showing everything
slurmtop --no-color # plain text
slurmtop --no-splash # skip the boot animation
slurmtop --qr # print only the QR code and exit
slurmtop --no-qr # hide the QR panel in the dashboard (on by default)
slurmtop --qr-wide # double-width QR modules, for fonts with gappy blocks
slurmtop --ascii # ASCII bars, for fonts without block glyphs
slurmtop --lang zh # ็น้ซไธญๆไป้ข๏ผ้ ่จญไพ $LANG ่ชๅๅคๆท๏ผ
slurmtop --title "lab-gpu" # header title (default: Slurm ClusterName)Run it on any node in the cluster. The node you are on is read locally; the rest are read over SSH, one round trip each per refresh.
| Element | Meaning |
|---|---|
โโโโโโโโ |
gradient bar, green โ yellow โ red |
โโโโ
โ |
sparkline of the last 24 refreshes โ tells idle-but-spiky apart from steadily pegged |
โฒ โผ |
trend against the last few samples |
| arc gauge | cluster GPU utilisation โ the dome fills left to right and is tinted by the value, so both shape and colour carry the reading |
โค โฅ โค โ |
HUD chrome โ section labels and frame ticks |
โ |
per-node status LED, tinted by that node's load |
| tinted panel background | the whole node zone warms up with its load โ amber past 45 %, orange past 75 %, pulsing red past 90 % โ so the node that is cooking is obvious without reading a single number |
| moving bright cell in a bar | scan sweep, advances every refresh |
โ โ ~ next to a GPU |
heat plume โ the GPU is โฅ70 ยฐC or โฅ95 % utilised |
| breathing bars | anything pegged at โฅ95 % pulses; so does a job within 15 % of its time limit |
โ FULL LOAD |
cluster mean utilisation โฅ90 %, blinking |
โฒ THERMAL |
hottest GPU โฅ78 ยฐC, blinking |
| LOAD panel | cluster utilisation on a sweeping scope โ data is written in a circle like an EKG, the bright column is the write head, and each cell uses eighth-blocks so six rows resolve 48 levels |
| bright cell running along a border | signal trace, one per panel at different phases |
GPUs โโโโโโโโ โ โโโโโโโโ |
one cell per GPU in the cluster, grouped by node โ the whole fleet at a glance |
โ run / โ pend |
Slurm job state; the running marker spins on every refresh |
PROG โโโโโโโโ |
how much of the job's time limit is used up โ turns red as it approaches the wall |
2h45m, 20m00s, 3d02h |
durations, always with units |
| header line | cluster totals: mean GPU utilisation, busy GPU count, VRAM, power draw, hottest GPU |
| panel border | tinted by that node's average GPU load |
By default nothing is hidden: every GPU row and every queued job is printed, even if the result is taller than the window. If the queue is long, the job list splits into two or three columns to claw back some height, but jobs are never dropped.
The layout fills the terminal: node panels split the full width evenly rather than sitting at a fixed size with dead space to the right, the queue takes whatever is left beside the QR panel, and spare vertical space goes to the LOAD scope. Sparklines need panels at least 58 columns wide.
If you would rather have a single screen that never scrolls, use --fit. That
mode gives up detail in order โ sparklines, then per-GPU rows collapsed to one
line per node, then multi-column jobs, and finally trimming the job list with
a N more job(s) hidden note.
If bars and sparklines show up as blank boxes or oddly wide blocks, your font
lacks the Unicode block glyphs. Use --ascii:
hipac-team3 [#...............] 6.2% .......... 1/16 GPUs ยท 129/2246G ยท 2.16 kW ยท 49ยฐC
โ CPU [..............] 0.6% load 5 6 7 โโ GPU0 [############] 99% 128.6G 49ยฐ 502W โ
The live view runs in the terminal's alternate screen buffer, so quitting restores whatever was on screen before and leaves no stack of stale frames in your scrollback.
- Reads only
nvidia-smi,/proc/stat,/proc/loadavg,free,squeueandsinfo. Nothing is written anywhere and no daemon is installed. - Sparkline history lives in the process, so it starts empty on each launch.
- AMD/Intel GPUs are not supported (patches welcome โ the only coupling is the
nvidia-smi --query-gpucall inREMOTE).
The dashboard carries a SCAN panel on the right with a scannable QR code for
this repo โ handy for getting the link onto someone's phone at a competition
without reading a URL out loud.
Each module is one character wide and half a character tall โ the upper half
of a cell is the foreground, the lower half the background โ so modules come
out square on any terminal whose cell is roughly 1:2. Verified by decoding
rendered output at cell ratios from 1.8 to 2.4, standalone and inside a full
dashboard frame. The panel is 44 columns and appears when at least 72 are left
for the queue; --no-qr hides it, --qr prints just the code.
If your font draws block characters with gaps and a scanner struggles,
--qr-wide redraws each module two characters wide using background colour
only, which does not depend on glyph coverage at all.
To print just the code:
slurmtop --qrThe matrix is embedded in the script, so no QR library is needed on the machine.
Built by team-03 / Hawks during HiPAC 2026 at NCHC. Issues and pull requests welcome.
MIT


