Skip to content

Latest commit

 

History

History
277 lines (215 loc) · 13.6 KB

File metadata and controls

277 lines (215 loc) · 13.6 KB

Deployment

How the site is deployed, how to change it, how to reload the data, and what it costs. The manifests themselves are documented in deploy/k8s/README.md; this page is everything around them.

What is where

Where What
deucepoint.net Cloudflare Pages The SPA, built from web/ on every push to main
api.deucepoint.net One VPS, k3s The API, Redis and Postgres, from deploy/k8s/
Images GHCR, public api, tools, web, built by CI on main, pinned by digest
DNS and TLS Cloudflare DNS; Let's Encrypt via cert-manager api is an A record, DNS-only, so the certificate is issued on the node

The host is a GreenCloud BudgetKVM in Staten Island: 4 EPYC Rome cores, 8 GB, 60 GB NVMe, 8 TB of transfer, Ubuntu 24.04, k3s v1.36. It replaced the Hetzner CX33 the Phase 4 issues priced, same shape at 40% of the cost; ADR-0008 and ADR-0009 carry the amendment.

The topology is the one ADR-0009 chose: everything in one node's cluster, Postgres on the node's own disk through local-path, the frontend deliberately outside the cluster on a CDN.

The host, as it was set up

Done once, by hand, on 14 September 2026. Not in the manifests because k3s and cert-manager are what the manifests assume rather than what they install.

# swap off for the kubelet; the 1 GB the panel created stays on disk, unused
swapoff -a && sed -i 's|^/swap.img|#/swap.img|' /etc/fstab
timedatectl set-timezone UTC

curl -sfL https://get.k3s.io | sh -                       # brings Traefik and local-path
kubectl apply -f https://github.com/cert-manager/cert-manager/releases/download/v1.21.2/cert-manager.yaml

ufw default deny incoming && ufw default allow outgoing
ufw allow 22/tcp && ufw allow 80/tcp && ufw allow 443/tcp
ufw allow from 10.42.0.0/16 && ufw allow from 10.43.0.0/16   # pods and services
ufw --force enable

SSH is key-only; the panel installed the key and disabled password login. The node's hostname is api.deucepoint.net, which is also the node name k3s reports.

Deploying the API

CI builds and pushes an image for every push to main that touches Go, the migrations or the Dockerfiles. Nothing deploys it. Deploying is:

  1. Take the digest from the CI log or the registry (deploy/k8s/README.md says how).
  2. Put it in base/40-api.yaml, and in the Jobs and base/60-load-weekly.yaml if tools moved. Commit.
  3. Copy and apply:
scp -r deploy/k8s deucepoint:/root/deploy/
ssh deucepoint kubectl apply -f /root/deploy/k8s/base/
ssh deucepoint kubectl -n deucepoint rollout status deploy/api

The Deployment rolls one pod at a time behind readiness, so the API stays up through it. Rolling back is the previous digest and the same three commands.

If the change carries a migration, run the migrate Job first, then apply. The Job's kubectl wait is in the README; an unmigrated database makes the new pods fail readiness and the rollout stalls rather than serving errors, which is the intended failure.

If the change adds a derived column or a derived table -- migration 00018 added both, sets and games from every score and player_totals -- run the refresh Job after the migration, once. It is the last step of every full load, so a reload does it anyway; on its own it takes a few minutes and the leaderboards and the year-by-year read as empty until it has run:

kubectl -n deucepoint delete job refresh --ignore-not-found
kubectl -n deucepoint apply -f /root/deploy/k8s/jobs/refresh.yaml
kubectl -n deucepoint logs -f job/refresh

This week (ADR-0016) is the ongoing-hourly CronJob and migration 00020. The first time: run the migrate Job, pin the tools image built with cmd/ongoing in base/65-ongoing-hourly.yaml, set suspend: false there, and apply. It refreshes at seven past every hour; to fill the card at once:

kubectl -n deucepoint create job --from=cronjob/ongoing-hourly ongoing-now
kubectl -n deucepoint logs -f job/ongoing-now

Deploying the frontend

Nothing to do. Cloudflare Pages watches main, runs npm run build in web/ with VITE_API_BASE=https://api.deucepoint.net/api/v1, and publishes dist/. Every pull request gets a preview URL on the CI checks. A bad build does not replace the live one.

Two things live only in the Pages project settings and not in the repository: that variable, and the custom domain.

Reloading the data

The database is reproducible from public files, and that is the recovery plan (#102): there is no backup to restore, because a rebuild produces the current schema from the current sources in about an hour and a restore would produce an old one.

The plan has been run end to end once, because standing the site up was the plan: on 14 September 2026 an empty volume took migrate in 4 seconds and load in 39 minutes, and smoke.sh passed against the result. Rebuilding on a new node is that plus the host setup above, and the DNS record.

What a rebuild does not bring back. Nothing, checked table by table. identity_reviews is the one that looked like it might: it is where the ingest queues an ambiguous match for a human, but the human's answer never goes into the database — it goes into configs/player_overrides.json, which is committed, and the next ingest re-queues whatever is still open. The ledgers (ingest_runs, ingest_files) are rewritten by the run that rebuilds them. Ratings, clutch and serve baselines are computed, never edited. The only thing on the node that is not derived from the repository or the sources is /root/secrets.yaml, the Postgres password, and a rebuild from empty can just mint a new one.

ssh deucepoint
kubectl -n deucepoint delete job load --ignore-not-found
kubectl -n deucepoint apply -f /root/deploy/k8s/jobs/load.yaml
kubectl -n deucepoint logs -f job/load -c ingest     # then -c rate

The Job is an init container running ingest, then a container running rate. Ingest is idempotent and resumable: every file's ETag is in ingest_files, an unchanged file costs a round trip, and a killed run picks up where it stopped. Ratings are recomputed from scratch every time and never patched, so there is nothing to carry over.

A partial refresh — the current season from the live source — is the same Job. The sources that changed are re-read; the 340 that did not are skipped.

This is what runs on a schedule. base/60-load-weekly.yaml is a CronJob of the same two containers at 06:00 UTC on Mondays, after the tour week it catches up on has finished, with concurrencyPolicy: Forbid so two can never overlap. Nothing is incremental, so a missed week costs nothing and the next run catches both. To see it, or to make it run now:

kubectl -n deucepoint get cronjob load-weekly
kubectl -n deucepoint create job --from=cronjob/load-weekly catchup-now
kubectl -n deucepoint logs -f job/catchup-now -c ingest

Failed runs are kept — three of them — and successful ones are not, so get jobs after a quiet week should be empty. Nothing alerts on a failure: the uptime check watches the site, and data a week behind is not an outage.

After a change to identity scoring or to configs/player_overrides.json — the reconcile stage on its own, then the ratings, which a merge invalidates:

sed 's/"reconcile"\]/"reconcile", "--dry-run"]/' deploy/k8s/jobs/reconcile.yaml   | kubectl -n deucepoint apply -f -              # what it would do, in the log
kubectl -n deucepoint logs -f job/reconcile -c reconcile
kubectl -n deucepoint delete job reconcile
kubectl -n deucepoint apply -f deploy/k8s/jobs/reconcile.yaml

The dry run is the review: a change to the scoring is judged on the pairs it would merge, listed one per line with the reason, before it merges any. Twenty seconds, then 1m 36s for the ratings.

After a change to configs/event_overrides.json, or to the rule that keys an event across seasons — the events stage on its own. Nothing the ratings read changes, so there is no rate step; about ten seconds:

kubectl -n deucepoint delete job events --ignore-not-found
kubectl -n deucepoint apply -f deploy/k8s/jobs/events.yaml
kubectl -n deucepoint logs -f job/events

Migration 00019 clears the name-keyed events for the stage to mint again, so a deploy that applies it runs this Job right after migrate; until it does, those events' pages answer 404. The log names every override that filed no row, which is how a renumbering a source has since undone gets noticed. A full load runs the same stage after the matches, so on an empty cluster this Job is never needed.

From empty (a new node, a lost volume): the README's first-time sequence — secrets, base/, migrate, load. The first migrate attempts fail while Postgres initialises its volume and the Job retries; that is expected.

Knowing it works

deploy/smoke.sh hits health, a player, a head-to-head and a simulation against a base URL and fails on the first non-200. Run it after every deploy and every load:

deploy/smoke.sh https://api.deucepoint.net

/api/v1/coverage is the claim the README rests on; check it after a load and make sure the dates are the ones the sources carry.

Knowing it is down

Nothing inside the cluster can report that the cluster is gone, so the watcher is outside it: the Uptime workflow (.github/workflows/uptime.yml) runs deploy/uptime.sh from a GitHub runner every fifteen minutes. It asks /api/v1/health for "database":"ok" and deucepoint.net for the page, retries once after thirty seconds so a blip is not a page, and fails the run otherwise. GitHub emails a failed scheduled run to whoever last committed the workflow file; that is the alert, and it reaches a phone.

Two things to know about it. A scheduled workflow is switched off after sixty days without a push to the repository, and GitHub says so by email when it does; the Run workflow button turns it back on. And the check is of the two public names, not the node — a check that passes says the site is up, and one that fails says only that it is not.

Runbook

Everything below starts from the laptop with bash deploy/kubectl.sh (on Windows, a bare .sh opens an editor), which is kubectl through an SSH tunnel to the node; no shell on the node is needed until a step says so.

The uptime run failed. deploy/uptime.sh https://api.deucepoint.net https://deucepoint.net from here says which of the two names it was. If it is the site, Cloudflare Pages has a status page and a deployment log, and nothing in this repository can fix it. If it is the API, deploy/kubectl.sh -n deucepoint get pods and read on. If kubectl cannot connect either, ssh deucepoint; if that cannot either, the GreenCloud panel has a console and a reboot button, and k3s starts on boot.

A pod is down. kubectl -n deucepoint get pods. The API self-heals behind readiness; Postgres is a StatefulSet and comes back on its volume; Redis comes back empty, which is allowed — every cache miss is a query. If a pod is Pending, the node is out of memory: kubectl describe node says which request did not fit.

Disk is full. df -h / on the node. Suspects in order: /var/lib/rancher/k3s (old images — k3s crictl rmi --prune), the Postgres volume under /var/lib/rancher/k3s/storage (a load doubles the database for its duration; 2.5 GB steady), and journald.

An ingest stopped half way. Re-apply the load Job. It resumes from the ledger; nothing half-written survives, because each batch is a transaction.

A stretch of matches is missing, and the source has them. The ledger is the usual cause rather than the source: the ingest asks conditionally and skips a 304, so a file first read mid-season is never re-read and the rest of that season never arrives. It shows on the site as a rating line with a hole in it. Count what is actually held and find the empty months first --

SELECT to_char(date_trunc('month', m.played_on), 'YYYY-MM') AS month, count(*)
  FROM matches m JOIN tournaments t ON t.id = m.tournament_id
 WHERE t.tour = 'wta' AND t.tier = 'tour' AND m.played_on >= '2025-01-01'
 GROUP BY 1 ORDER BY 1;

-- then jobs/load-force.yaml, narrowed to the tour and seasons concerned as its header shows. If the months are still empty after a forced read, the source genuinely lacks them and /api/v1/coverage should be saying so.

The certificate is expiring. cert-manager renews thirty days out and writes to samisaleh07@gmail.com if it cannot. kubectl -n deucepoint get certificate shows Ready; kubectl -n deucepoint get challenge shows what is stuck. The HTTP-01 challenge needs port 80 open and api.deucepoint.net resolving to the node, DNS-only, not proxied.

Logs. deploy/kubectl.sh -n deucepoint logs deploy/api --since=1h, and -c ingest or -c rate on job/load. The first run copies the cluster's kubeconfig to ~/.kube/deucepoint.yaml — it is the admin credential, so it stays there — and opens the tunnel, which stays up until the ssh process is killed or the laptop sleeps; the script reopens it. The kubelet keeps five files of 10 MB per container and nothing keeps a finished pod's, so the logs of the pod a rollout replaced are gone with it; nothing ships them anywhere, and at this traffic nothing needs to.

What it costs

VPS GreenCloud BudgetKVM NYC, annual $45.00 / year
Domain deucepoint.net, Cloudflare Registrar, at cost $11.86 / year
Pages, GHCR, DNS, certificates 0

$56.86 a year, $4.74 a month. The Phase 4 issues estimated about 9 EUR a month on Hetzner; this is half that. No part of it scales with visitors. The VPS is prepaid and non-refundable; if the project ended tomorrow, the year is the loss, and the box is a general-purpose 8 GB machine until then.