Repository navigation
fix(workspaces): read hosts concurrently so the page stops hanging - #2017
Conversation
The Workspaces page opens on the flat "every file" listing, which read up to 25 container-backed workspaces one after another, then fetched their thumbnails one after another, each on a fresh archive client. The page waited for the sum of those round trips, and a host that had gone away held it for the archive's 60-second default timeout. Workspaces and thumbnails are now read concurrently, at most eight at a time, and a listing gives up on a host after ten seconds. The thumbnail budget is still shared out in listing order before anything is fetched, so which tiles are drawn does not depend on which host answers first. Connection resolution is serialised behind a lock because the concurrent readers share one database session. The "Count files" switch reads hosts the same way.
Review follow-up on the concurrent reads. The fan-outs run in a TaskGroup, so a call that raises cancels the rest instead of leaving them running against the request's session. Listings and thumbnails share one semaphore per request, taken around single host calls only, so one page load has at most eight executor threads in flight rather than eight per host. Adds tests for the connection lock and the per-request bound, makes the budget-ordering test land the first host last with an Event, and exempts it from the security marker: it is a fetch budget, not spend. The archive always takes a timeout, defaulting to the library's. Refs #2008
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
Security findingsFinding details are still loading. Check the individual review comments. ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 318ea20224
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…rent-reads # Conflicts: # CHANGELOG.md
…rent-reads # Conflicts: # CHANGELOG.md
The per-request semaphore bounded one page load only: the service is built per request, so a few concurrent loads against a dead host could take every thread of the loop's default executor, which bcrypt and DNS also run on. Host ls and read_bytes calls now run on a process-wide eight-thread pool, carrying the caller's context as to_thread does, and the per-request semaphore goes. The test runs two page loads at once and fails on the per-request version with 14 calls in flight.
### Summary Release 0.0.522: version, lock and changelog. ### Changes - `backend/pyproject.toml`, `frontend/package.json` and `backend/uv.lock` move to 0.0.522. - `CHANGELOG.md`: the `[Unreleased]` block becomes `[0.0.522] - 2026-10-06`, with a fresh empty `[Unreleased]` above it. - Ships #2017 (from #2008): the Workspaces page reads hosts and thumbnails side by side on a process-wide pool of eight threads of its own, and a listing gives up on an unanswered host call after ten seconds.
Summary
Replaces #2008 by @henfircreo: the Workspaces page reads hosts concurrently instead of one after another. This branch holds only that commit, cherry-picked onto
mainwith its authorship kept, plus the follow-ups the review on #2008 asked for. The fork's other changes (refresh-token reissue, 1800 s timeout ceiling, Dockerfile, SQL script) are not here; the auth one is #2007.Changes
250d254c8(henfircreo): listings and thumbnails read side by side, the thumbnail budget shared out in listing order before any fetch, a 10 s browse timeout per host call, connection resolution behind a lock.asyncio.TaskGroup, so a call that raises cancels the rest instead of leaving them running against the request's session.lsandread_bytescalls run on a process-wide eight-threadThreadPoolExecutorof their own (_HOST_CALLS), not the default executorto_threaduses, whichbcryptand DNS share. A per-request semaphore was tried first; Codex pointed out it bounds one page load only, so a few concurrent loads against a dead host could still take every default thread. The pool bounds the process, and a dead host slows only other host calls._archivealways passes a timeout, defaulting to the library'sDEFAULT_TIMEOUT_SECONDS.Eventso the first-listed host really answers last, and it is exempt from the security marker (fetch budget, not spend).Verification
tests/test_sandbox_workspace.py,tests/api/test_workspace_browser_routes.py,tests/test_security_marker.py: 239 passed.app/services/sandbox_workspace.pyat 100% coverage.ruff format,ruff check,ty checkon the changed files: clean.make testandmake check; CI covers those.Closes #2008