Skip to content

chore: refresh repository metadata - #12

Closed
rayenmabrouk wants to merge 16 commits into
masterfrom
chore/refresh-metadata
Closed

rayenmabrouk wants to merge 16 commits into
masterfrom
chore/refresh-metadata

Conversation

@rayenmabrouk

Copy link
Copy Markdown
Owner

No description provided.

Original dpaste README moved to docs/upstream-dpaste-README.md. README separates upstream application code from the infrastructure work, lists verified behaviour, trade-offs and AI-assistance disclosure. Screenshots renamed chronologically; admin IP redacted.
…, 5xx alarm

- default_tags (Project, Environment, ManagedBy, Repository) on every resource
- Environment-specific values exposed as validated root variables (instance type,
  key pair, instance profile, log retention, image retention, alarm email)
- SSH ingress is now opt-in: empty allowed_ssh_cidr closes port 22 (ops use SSM);
  0.0.0.0/0 is rejected by validation
- ECR lifecycle policy: keep the 10 newest images, expire untagged after 1 day
- S3 backup bucket for SQLite (versioned, SSE-S3, TLS-only, public access
  blocked, 14-day expiry) + SSM parameter with its name
- CloudWatch: log metric filter on Caddy access logs + HTTP 5xx alarm (catches
  a dead dpaste container, which the EC2 status check cannot see); optional SNS
  email notifications for all alarms
- Modules declare required_providers; unused variables removed (tflint clean)
- More outputs: instance_id, SSM session command, backup bucket, alarm names
- Checkov: 47 passed, 0 failed, 17 skipped with inline justification
- scripts/backup.sh: online SQLite backup (sqlite3 backup API) to the S3 backup
  bucket, list, and restore with integrity check + pre-restore backup
- deploy.sh installs a daily backup timer (non-fatal if backup.sh is missing),
  reads the region from IMDSv2 instead of hard-coding it, and removes old release
  images from the instance disk while keeping the rollback target
… scan

CI
- All third-party actions pinned to full commit SHAs (trivy-action was @master)
- permissions: contents: read; concurrency cancels superseded PR runs; timeouts
- lint is now blocking (ruff.toml documents the one ignored upstream rule)
- docker-build-scan: hadolint (checksum-verified binary) -> build -> container
  smoke test (health check healthy, non-root uid, GET / and API POST) -> Trivy
- new secret-scan job: gitleaks over the full history (.gitleaksignore lists
  one 2013 upstream Coveralls token)

CD
- Image is built and scanned before AWS credentials are configured, so the
  build and the third-party scanner never run with cloud access
- Ships backup.sh with deploy.sh (gzip + base64: smaller SSM payload than before)

Terraform pipeline: actions pinned, TFLint step added

Dockerfile: node:24-slim instead of floating node:lts-slim, PYTHONUNBUFFERED
for prompt CloudWatch logs, OCI labels, exec-form HEALTHCHECK, quoted extras

Repository hygiene
- Dependabot for github-actions, docker, terraform and pip; renovate.json
  removed (two bots would open duplicate PRs)
- Removed upstream leftovers that no longer apply: .travis.yml, .deploystack/,
  docker.yml.disabled, minimal.docker-compose.yml, original root-running
  Dockerfile, tox.ini (referenced missing README.rst), issue templates,
  CONTRIBUTING.md, CODE_OF_CONDUCT.md
- .python-version/.node-version match the image (3.10 / 24)
- .trivyignore BOM removed; .gitignore covers .env, *.pem, plan files
- monitoring stack mounts ~/.aws on Linux/macOS as well as Windows
README restructured: architecture diagram (only resources that exist),
inherited vs built, infrastructure with the reason for each component, alarms
and what they detect, security controls and the least-privilege IAM a real
account would use, CI/CD, exact local/AWS commands, verification table,
cost, AWS Academy limitations, project structure, engineering decisions.
Clearly separates what was verified in AWS from what was added afterwards.
Removes the links to docs/architecture.md and docs/security.md, which did not
exist; uses only measured image sizes.

Runbook: SSH now optional, backups/restore, alarms and SNS, Session Manager,
teardown of the backup bucket, list-nested code blocks fixed (the PowerShell
here-string terminator was indented and failed when pasted). New or changed
steps are marked [not yet verified].

Cost: third alarm, custom metric, backup bucket, SNS; placeholder row removed.
"text" is rejected by the API with HTTP 400; verified _text returns 200.
The Academy service control policy denies s3:GetBucketObjectLockConfiguration,
which the AWS provider calls whenever it reads an aws_s3_bucket resource, so
the Terraform-managed backup bucket failed on every plan after creation.

- The bucket is now created by scripts/bootstrap-tfstate.ps1 (idempotent),
  like the state bucket
- Terraform keeps managing its public access block, versioning, encryption,
  lifecycle and TLS-only policy; none of these read object lock
- Docs: limitation, runbook bootstrap/teardown, troubleshooting #16 (stale
  saved plan) and #17 (this SCP)

Existing state must drop the old resource first (the bucket stays in AWS):
  terraform state rm module.compute.aws_s3_bucket.backups
Bootstrap the backup bucket outside Terraform (AWS Academy SCP)
Pipeline checks, post-restart redeploy, infrastructure apply without
replacement or drift, and an on-demand S3 backup are now marked verified.
Restore, 5xx alarm firing, SNS and the scheduled backup remain unverified.
Docs: record what was verified in AWS on 2026-09-24
docs: remove 'How this was built' section from README
@rayenmabrouk
rayenmabrouk deleted the chore/refresh-metadata branch September 26, 2026 10:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant