How the site is deployed, how to change it, how to reload the data, and what it costs.
The manifests themselves are documented in deploy/k8s/README.md;
this page is everything around them.
| Where | What | |
|---|---|---|
deucepoint.net |
Cloudflare Pages | The SPA, built from web/ on every push to main |
api.deucepoint.net |
One VPS, k3s | The API, Redis and Postgres, from deploy/k8s/ |
| Images | GHCR, public | api, tools, web, built by CI on main, pinned by digest |
| DNS and TLS | Cloudflare DNS; Let's Encrypt via cert-manager | api is an A record, DNS-only, so the certificate is issued on the node |
The host is a GreenCloud BudgetKVM in Staten Island: 4 EPYC Rome cores, 8 GB, 60 GB NVMe, 8 TB of transfer, Ubuntu 24.04, k3s v1.36. It replaced the Hetzner CX33 the Phase 4 issues priced, same shape at 40% of the cost; ADR-0008 and ADR-0009 carry the amendment.
The topology is the one ADR-0009 chose: everything in one node's cluster, Postgres on the
node's own disk through local-path, the frontend deliberately outside the cluster on a CDN.
Done once, by hand, on 14 September 2026. Not in the manifests because k3s and cert-manager are what the manifests assume rather than what they install.
# swap off for the kubelet; the 1 GB the panel created stays on disk, unused
swapoff -a && sed -i 's|^/swap.img|#/swap.img|' /etc/fstab
timedatectl set-timezone UTC
curl -sfL https://get.k3s.io | sh - # brings Traefik and local-path
kubectl apply -f https://github.com/cert-manager/cert-manager/releases/download/v1.21.2/cert-manager.yaml
ufw default deny incoming && ufw default allow outgoing
ufw allow 22/tcp && ufw allow 80/tcp && ufw allow 443/tcp
ufw allow from 10.42.0.0/16 && ufw allow from 10.43.0.0/16 # pods and services
ufw --force enableSSH is key-only; the panel installed the key and disabled password login. The node's
hostname is api.deucepoint.net, which is also the node name k3s reports.
CI builds and pushes an image for every push to main that touches Go, the migrations or
the Dockerfiles. Nothing deploys it. Deploying is:
- Take the digest from the CI log or the registry (
deploy/k8s/README.mdsays how). - Put it in
base/40-api.yaml, and in the Jobs andbase/60-load-weekly.yamliftoolsmoved. Commit. - Copy and apply:
scp -r deploy/k8s deucepoint:/root/deploy/
ssh deucepoint kubectl apply -f /root/deploy/k8s/base/
ssh deucepoint kubectl -n deucepoint rollout status deploy/apiThe Deployment rolls one pod at a time behind readiness, so the API stays up through it. Rolling back is the previous digest and the same three commands.
If the change carries a migration, run the migrate Job first, then apply. The Job's
kubectl wait is in the README; an unmigrated database makes the new pods fail readiness
and the rollout stalls rather than serving errors, which is the intended failure.
If the change adds a derived column or a derived table -- migration 00018 added both,
sets and games from every score and player_totals -- run the refresh Job after the
migration, once. It is the last step of every full load, so a reload does it anyway; on
its own it takes a few minutes and the leaderboards and the year-by-year read as empty
until it has run:
kubectl -n deucepoint delete job refresh --ignore-not-found
kubectl -n deucepoint apply -f /root/deploy/k8s/jobs/refresh.yaml
kubectl -n deucepoint logs -f job/refreshThis week (ADR-0016) is the ongoing-hourly CronJob and migration 00020. The first
time: run the migrate Job, pin the tools image built with cmd/ongoing in
base/65-ongoing-hourly.yaml, set suspend: false there, and apply. It refreshes at seven
past every hour; to fill the card at once:
kubectl -n deucepoint create job --from=cronjob/ongoing-hourly ongoing-now
kubectl -n deucepoint logs -f job/ongoing-nowNothing to do. Cloudflare Pages watches main, runs npm run build in web/ with
VITE_API_BASE=https://api.deucepoint.net/api/v1, and publishes dist/. Every pull
request gets a preview URL on the CI checks. A bad build does not replace the live one.
Two things live only in the Pages project settings and not in the repository: that variable, and the custom domain.
The database is reproducible from public files, and that is the recovery plan (#102): there is no backup to restore, because a rebuild produces the current schema from the current sources in about an hour and a restore would produce an old one.
The plan has been run end to end once, because standing the site up was the plan: on
14 September 2026 an empty volume took migrate in 4 seconds and load in 39 minutes,
and smoke.sh passed against the result. Rebuilding on a new node is that plus the host
setup above, and the DNS record.
What a rebuild does not bring back. Nothing, checked table by table. identity_reviews
is the one that looked like it might: it is where the ingest queues an ambiguous match for
a human, but the human's answer never goes into the database — it goes into
configs/player_overrides.json, which is committed, and the next ingest re-queues whatever
is still open. The ledgers (ingest_runs, ingest_files) are rewritten by the run that
rebuilds them. Ratings, clutch and serve baselines are computed, never edited. The only
thing on the node that is not derived from the repository or the sources is
/root/secrets.yaml, the Postgres password, and a rebuild from empty can just mint a new
one.
ssh deucepoint
kubectl -n deucepoint delete job load --ignore-not-found
kubectl -n deucepoint apply -f /root/deploy/k8s/jobs/load.yaml
kubectl -n deucepoint logs -f job/load -c ingest # then -c rateThe Job is an init container running ingest, then a container running rate. Ingest is
idempotent and resumable: every file's ETag is in ingest_files, an unchanged file costs a
round trip, and a killed run picks up where it stopped. Ratings are recomputed from
scratch every time and never patched, so there is nothing to carry over.
A partial refresh — the current season from the live source — is the same Job. The sources that changed are re-read; the 340 that did not are skipped.
This is what runs on a schedule. base/60-load-weekly.yaml is a CronJob of the same two
containers at 06:00 UTC on Mondays, after the tour week it catches up on has finished,
with concurrencyPolicy: Forbid so two can never overlap. Nothing is incremental, so a
missed week costs nothing and the next run catches both. To see it, or to make it run
now:
kubectl -n deucepoint get cronjob load-weekly
kubectl -n deucepoint create job --from=cronjob/load-weekly catchup-now
kubectl -n deucepoint logs -f job/catchup-now -c ingestFailed runs are kept — three of them — and successful ones are not, so get jobs after a
quiet week should be empty. Nothing alerts on a failure: the uptime check watches the
site, and data a week behind is not an outage.
After a change to identity scoring or to configs/player_overrides.json — the
reconcile stage on its own, then the ratings, which a merge invalidates:
sed 's/"reconcile"\]/"reconcile", "--dry-run"]/' deploy/k8s/jobs/reconcile.yaml | kubectl -n deucepoint apply -f - # what it would do, in the log
kubectl -n deucepoint logs -f job/reconcile -c reconcile
kubectl -n deucepoint delete job reconcile
kubectl -n deucepoint apply -f deploy/k8s/jobs/reconcile.yamlThe dry run is the review: a change to the scoring is judged on the pairs it would merge, listed one per line with the reason, before it merges any. Twenty seconds, then 1m 36s for the ratings.
After a change to configs/event_overrides.json, or to the rule that keys an event
across seasons — the events stage on its own. Nothing the ratings read changes, so
there is no rate step; about ten seconds:
kubectl -n deucepoint delete job events --ignore-not-found
kubectl -n deucepoint apply -f deploy/k8s/jobs/events.yaml
kubectl -n deucepoint logs -f job/eventsMigration 00019 clears the name-keyed events for the stage to mint again, so a deploy
that applies it runs this Job right after migrate; until it does, those events' pages
answer 404. The log names every override that filed no row, which is how a renumbering
a source has since undone gets noticed. A full load runs the same stage after the matches, so on an
empty cluster this Job is never needed.
From empty (a new node, a lost volume): the README's first-time sequence — secrets,
base/, migrate, load. The first migrate attempts fail while Postgres initialises its
volume and the Job retries; that is expected.
deploy/smoke.sh hits health, a player, a head-to-head and a simulation against a base
URL and fails on the first non-200. Run it after every deploy and every load:
deploy/smoke.sh https://api.deucepoint.net/api/v1/coverage is the claim the README rests on; check it after a load and make sure
the dates are the ones the sources carry.
Nothing inside the cluster can report that the cluster is gone, so the watcher is outside
it: the Uptime workflow (.github/workflows/uptime.yml) runs deploy/uptime.sh from a
GitHub runner every fifteen minutes. It asks /api/v1/health for "database":"ok" and
deucepoint.net for the page, retries once after thirty seconds so a blip is not a page,
and fails the run otherwise. GitHub emails a failed scheduled run to whoever last
committed the workflow file; that is the alert, and it reaches a phone.
Two things to know about it. A scheduled workflow is switched off after sixty days without
a push to the repository, and GitHub says so by email when it does; the Run workflow
button turns it back on. And the check is of the two public names, not the node — a check
that passes says the site is up, and one that fails says only that it is not.
Everything below starts from the laptop with bash deploy/kubectl.sh (on Windows, a bare
.sh opens an editor), which is kubectl
through an SSH tunnel to the node; no shell on the node is needed until a step says so.
The uptime run failed. deploy/uptime.sh https://api.deucepoint.net https://deucepoint.net
from here says which of the two names it was. If it is the site, Cloudflare Pages has a
status page and a deployment log, and nothing in this repository can fix it. If it is the
API, deploy/kubectl.sh -n deucepoint get pods and read on. If kubectl cannot connect
either, ssh deucepoint; if that cannot either, the GreenCloud panel has a console and a
reboot button, and k3s starts on boot.
A pod is down. kubectl -n deucepoint get pods. The API self-heals behind readiness;
Postgres is a StatefulSet and comes back on its volume; Redis comes back empty, which is
allowed — every cache miss is a query. If a pod is Pending, the node is out of memory:
kubectl describe node says which request did not fit.
Disk is full. df -h / on the node. Suspects in order: /var/lib/rancher/k3s (old
images — k3s crictl rmi --prune), the Postgres volume under /var/lib/rancher/k3s/storage
(a load doubles the database for its duration; 2.5 GB steady), and journald.
An ingest stopped half way. Re-apply the load Job. It resumes from the ledger; nothing half-written survives, because each batch is a transaction.
A stretch of matches is missing, and the source has them. The ledger is the usual cause rather than the source: the ingest asks conditionally and skips a 304, so a file first read mid-season is never re-read and the rest of that season never arrives. It shows on the site as a rating line with a hole in it. Count what is actually held and find the empty months first --
SELECT to_char(date_trunc('month', m.played_on), 'YYYY-MM') AS month, count(*)
FROM matches m JOIN tournaments t ON t.id = m.tournament_id
WHERE t.tour = 'wta' AND t.tier = 'tour' AND m.played_on >= '2025-01-01'
GROUP BY 1 ORDER BY 1;-- then jobs/load-force.yaml, narrowed to the tour and seasons concerned as its header
shows. If the months are still empty after a forced read, the source genuinely lacks them
and /api/v1/coverage should be saying so.
The certificate is expiring. cert-manager renews thirty days out and writes to
samisaleh07@gmail.com if it cannot. kubectl -n deucepoint get certificate shows Ready;
kubectl -n deucepoint get challenge shows what is stuck. The HTTP-01 challenge needs port
80 open and api.deucepoint.net resolving to the node, DNS-only, not proxied.
Logs. deploy/kubectl.sh -n deucepoint logs deploy/api --since=1h, and -c ingest
or -c rate on job/load. The first run copies the cluster's kubeconfig to
~/.kube/deucepoint.yaml — it is the admin credential, so it stays there — and opens the
tunnel, which stays up until the ssh process is killed or the laptop sleeps; the script
reopens it. The kubelet keeps five files of 10 MB per container and nothing keeps a
finished pod's, so the logs of the pod a rollout replaced are gone with it; nothing ships
them anywhere, and at this traffic nothing needs to.
| VPS | GreenCloud BudgetKVM NYC, annual | $45.00 / year |
| Domain | deucepoint.net, Cloudflare Registrar, at cost |
$11.86 / year |
| Pages, GHCR, DNS, certificates | 0 |
$56.86 a year, $4.74 a month. The Phase 4 issues estimated about 9 EUR a month on Hetzner; this is half that. No part of it scales with visitors. The VPS is prepaid and non-refundable; if the project ended tomorrow, the year is the loss, and the box is a general-purpose 8 GB machine until then.