Repository navigation
fix(db): bound connection pools and back off edge heartbeats - #486
Merged
Merged
Conversation
Limit MySQL connections per Manager, retain idle connections, and expose pool budgets in both Compose stacks. Back off failed heartbeats without immediate registration writes or process restarts while preserving tunnel recovery and upgrade health markers. Refs #473
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Concurrent device heartbeats could open an unbounded number of MySQL connections. Each failed heartbeat then triggered registration writes, and repeated failures could restart the Edge process, adding more load during a database outage.
Refs #473. This addresses the verified connection-pressure amplification mechanism; the initiating customer trigger and the contribution of the slow alert query remain unconfirmed.
Validation
go test -raceandgo vetfor config, dbx, Edge biz, and tunnel. Re-ran Edge race/vet after the final reconnect health-marker change.go test -p 2 ./...in an isolated source copy as a non-root user. macOS cannot run all Linux-only packages; initial Linux container failures from a stale read-only module cache, a host worktree path, and root permission semantics were resolved by correcting the test environment.ONGRID_TEST_DB_POOL=1 TESTCONTAINERS_RYUK_DISABLED=true go test -race -tags=integration ./tests/integration -run '^TestMySQLPoolHeartbeatBurst$' -count=1 -v: 220 devices, 660 heartbeats, zero heartbeat failures and zero MySQL connection-limit rejections at servermax_connections=151. Verified pool saturation honors request deadlines, releases connections, recovers, and accepts a custom limit. The test container was removed.go test -tags=e2e ./tests/e2e -run '^TestAuth_LoginAndSelf_B1$' -count=1with a disposable MySQL and actual Manager process.git diff --checkpassed.make arch-lintcould not run becausego-arch-lintis not installed. No customer-environment validation or release has been performed.Risk and rollback
The total budget must cover every Manager replica and other database clients, with operational headroom. An undersized pool can increase queueing/timeouts; slow queries or database resource problems still require investigation. Upgrade Manager first, then Edge; old Edge versions retain the immediate-registration retry behavior.
No schema, API, or dependency changes. Adjust pool environment variables and recreate Manager to change the budget. Reverting the images restores the old unbounded pool/retry behavior and therefore the original failure risk. See
docs/install/mysql-connection-pool.mdfor configuration and acceptance steps.Author confirmation