Skip to content

Monitoring: Prometheus + Grafana (black-box probes + CloudWatch) - #7

Merged
rayenmabrouk merged 1 commit into
masterfrom
feature/monitoring-stack
Sep 23, 2026
Merged

rayenmabrouk merged 1 commit into
masterfrom
feature/monitoring-stack

Conversation

@rayenmabrouk

Copy link
Copy Markdown
Owner

Summary

Local observability stack (monitoring/, Docker Compose) that monitors both the local container and the live AWS deployment, as black-box monitoring: no changes to the dpaste application code.

Components (versions pinned, ports bound to 127.0.0.1)

  • blackbox_exporter v0.25.0: HTTP probes of http://dpaste:8000/ (local) and the production HTTPS URL. Availability, total latency, HTTP status, TLS certificate expiry, per-phase timing (DNS, connect, TLS, processing, transfer)
  • Prometheus v3.2.1: 15s scrape, 7-day retention, alert rules:
    • DpasteDown: probe failing for 1 minute (critical)
    • DpasteSlowResponse: probe > 2s for 5 minutes (warning)
    • TLSCertificateExpiringSoon: certificate < 14 days (warning)
  • Grafana 11.6.0: datasources and dashboard provisioned from files (dashboards as code)
    • Prometheus: probe metrics and alert count
    • CloudWatch: production EC2 CPU (1-minute detailed monitoring) and network, via read-only mounted AWS credentials

set-production-target.ps1 writes the production probe target from terraform output app_url (the sslip.io hostname changes with the EC2 IP). The generated file is git-ignored.

Verification

  • All Prometheus targets UP; probe_success = 1 for local and production
  • Alert test: stopped the local container, DpasteDown went pending then firing, Grafana showed Local DOWN and 1 firing alert; restarted, alert resolved
  • CloudWatch panels show production CPU/network, including the spikes from CD deployments

Decision: cAdvisor removed

cAdvisor v0.49.1 could not identify containers on Docker Desktop (containerd image store: failed to identify the read-write layer ID), so it exported no per-container series. Replaced by CloudWatch metrics for the production instance instead of shipping a component that does not work.

Limitations

  • No Alertmanager: alerts are visible in Prometheus and Grafana but not routed to email/Slack
  • CloudWatch panels use AWS Academy session credentials and stop working when the lab session ends

…loudWatch panels

Blackbox exporter probes the local container and the live AWS deployment (availability, latency, HTTP status, TLS expiry, request phases) without changing dpaste code. Prometheus evaluates DpasteDown, DpasteSlowResponse and TLSCertificateExpiringSoon. Grafana datasources and dashboard are provisioned from files; production EC2 CPU and network come from CloudWatch.

cAdvisor evaluated and removed: incompatible with Docker Desktop's containerd image store (cannot resolve container layer IDs).
@rayenmabrouk
rayenmabrouk merged commit aed629f into master Sep 23, 2026
3 checks passed
@rayenmabrouk
rayenmabrouk deleted the feature/monitoring-stack branch September 23, 2026 20:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant