Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -185,7 +185,9 @@ In a different AWS account, pass the state bucket at init time: `terraform init

**Verified in AWS** (executed and observed, not just configured): CD deploys through SSM; **automatic rollback** after a deliberately broken image (production kept serving 200); data persisted across container replacement; Terraform apply paused for approval and applied the exact plan; `DpasteDown` fired in Prometheus/Grafana when the container stopped and resolved after restart; CloudWatch received container logs.

**Added after that verification and not yet exercised in AWS:** SQLite backups to S3 and restore, the HTTP 5xx metric filter and alarm, SNS notifications, the ECR lifecycle policy, the optional SSH rule, and the updated CI/CD workflows. The runbook marks them `[not yet verified]`.
**Verified on 2026-09-24 after the hardening changes:** all CI checks green on the pull request (tests, lint, hadolint, image build + container smoke test, Trivy, gitleaks, Terraform checks and plan); CD redeployed after a lab restart moved the instance to a new IP, and the site answered HTTP 200 over HTTPS on the new hostname; the infrastructure changes (tags, ECR lifecycle policy, backup bucket settings, 5xx metric filter and alarm) were applied through the approval gate with no replacement, and the next plan showed no drift; an on-demand SQLite backup uploaded to the encrypted S3 bucket.

**Not yet exercised:** restoring a backup, the 5xx alarm firing, SNS email notifications, the SSH-disabled configuration, and the scheduled (03:00 UTC) backup run. The runbook marks them `[not yet verified]`.

## Cost considerations

Expand Down
8 changes: 4 additions & 4 deletions docs/deployment-runbook.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,7 @@ Academy credentials expire when the lab session ends (about 4 hours), and the EC

5. If the AWS Console shows `explicit deny ... voc-cancel-cred`, the console tab is from an older session: close all console tabs and reopen the console from Vocareum.

### After a lab restart: redeploy [not yet verified]
### After a lab restart: redeploy [verified 2026-09-24]

When the instance starts again it gets a **new public IP**, so the sslip.io hostname changes. The containers restart automatically, but Caddy's certificate and Django's `ALLOWED_HOSTS` still refer to the old hostname. Redeploy so `deploy.sh` regenerates both:

Expand Down Expand Up @@ -101,7 +101,7 @@ cd ..

This creates the VPC, subnet, internet gateway, security groups, EC2 instance (Docker installed by `user_data`), ECR repository with lifecycle policy, SSM `SecureString` parameter with a generated Django `SECRET_KEY`, the backup bucket's settings (versioning, encryption, lifecycle, TLS-only policy, public access block) and its SSM parameter, CloudWatch log group, 5xx metric filter and alarms (and an SNS topic if `alarm_email` is set).

**[not yet verified]** The ECR lifecycle policy, backup bucket, 5xx metric filter/alarm, SNS topic and the optional SSH rule were added after the verified build. The first plan after pulling these changes should show only additions (ECR lifecycle policy, S3 bucket and its settings, SSM parameter, metric filter, alarm) and in-place tag updates, **no replacement**. Stop and investigate if it proposes to replace the instance.
**[verified 2026-09-24]** The ECR lifecycle policy, backup bucket settings and 5xx metric filter/alarm were applied through the pipeline with in-place tag updates and **no replacement**; the next plan reported no changes. Always stop and investigate if a plan proposes to replace the instance. The SNS topic (`alarm_email`) is **[not yet verified]**.

After this first bootstrap, **infrastructure changes go through pull requests** and the Terraform pipeline (section 4.2).

Expand Down Expand Up @@ -232,7 +232,7 @@ aws ecr describe-images --repository-name cloudpulse --query "sort_by(imageDetai

Runs daily via `cloudpulse-cleanup.timer`. Run it now: `systemctl start cloudpulse-cleanup.service` (through SSM), then `journalctl -u cloudpulse-cleanup.service -n 5`.

### Backups and restore [not yet verified]
### Backups and restore [backup verified 2026-09-24; restore and scheduled run not yet verified]

`cloudpulse-backup.timer` runs `/opt/cloudpulse/backup.sh backup` daily at 03:00 UTC (and at boot if a run was missed while the lab was stopped). It uses SQLite's online backup API, so dpaste keeps serving, and uploads `sqlite/dpaste-<UTC timestamp>.sqlite.gz` to the backup bucket (kept 14 days).

Expand All @@ -247,7 +247,7 @@ journalctl -u cloudpulse-backup.service -n 20 # last run

`restore` downloads the backup, refuses it unless `PRAGMA integrity_check` returns `ok`, takes a fresh backup of the current database, stops dpaste, replaces the database file in the `dpaste_data` volume, restarts dpaste and waits for it to be healthy. From your machine: `aws s3 ls s3://$(terraform -chdir=terraform output -raw backup_bucket_name)/sqlite/`.

### Alarms [not yet verified for the 5xx alarm and SNS]
### Alarms [5xx alarm created and OK; firing and SNS not yet verified]

```powershell
aws cloudwatch describe-alarms --alarm-name-prefix cloudpulse --query "MetricAlarms[].[AlarmName,StateValue]" --output table
Expand Down
Loading