fix: lift the 30 s timeout on the 0.3.5 migration's data-sized steps - #160
Merged
Merged
Conversation
The 0.3.5 migration ran its recursive chown of the PostgreSQL cluster, the find -exec chmod over the app tree, and the per-directory chmod and listing of data/ through execFail with no timeout argument, so each inherited the SDK's 30 s default and was SIGKILLed on any instance where one of them took longer — a cold HDD over tens of thousands of app files, or a single huge data/ directory — failing the migration. They now run unbounded. 34.0.4:5 is not yet on beta, so the fix rides that version and joins its release note. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Four steps of the 0.3.5 migration run through
execFailwithout a timeout:chownof the PostgreSQL clusterfind -exec chmodover the app treechmodofdata/data/With no timeout, each gets the SDK's 30 s default and is SIGKILLed on any instance where it runs longer: a cold HDD working through tens of thousands of app files, or a single very large directory under
data/. When that happens the migration fails. All four now run with no timeout (null).34.0.4:5 isn't on beta yet, so this doesn't bump the version: the fix ships in 34.0.4:5 and is added to its release note. Draft #157 plans to move past whatever version
masterhas when it's finished, andnext(#153) goes to 35.0.0:0, so neither conflicts with this beyond the release-note line.The comment above the
data/walk says a single recursivefindorchmod -Rwas OOM-killed (SIGKILL). With this 30 s default in effect that SIGKILL may have been the timeout instead. I haven't changed the walk: it works either way, and memory use stays bounded.Found in the fleet-wide audit for this 30 s default after Start9Labs/open-webui-startos#43.
🤖 Generated with Claude Code