Conversation
filimonov
marked this pull request as draft
October 1, 2026 14:35
filimonov
force-pushed
the
fix/antalya-26.6/cas-gc-janitor-batches
branch
2 times, most recently
from
October 1, 2026 16:34
c2fef81 to
303a40c
Compare
The namespace janitor (GC phase 16) deleted dead `_log`/`_snap` keys one page per folding round: a dropped table with 20k parts took 16 rounds and 16k DELETE requests to drain. It now deletes a page's dead keys in batches sized by the store's batch-delete limit (`objects_chunk_size_to_delete` on S3, 1 elsewhere), runs the batches as jobs on the GC I/O pool (the pool's `cas_gc_io_concurrency` threads bound them, the same pattern as the pending-deletes fan-out) while the next page is listed, and keeps taking pages until a soft 20 s budget: no new page after it, in-flight jobs complete. A pass that began mid-stream wraps once to the stream start. Rules: - Dead `_log`/`_snap` keys go in write-once cohorts with no per-key precondition and no per-key fallback; a batch that still fails after retries is leaked and the cursor advances. - A page is held (cursor not advanced, nothing published) only on lost authority, a capability refusal (`NOT_IMPLEMENTED`), a schedule refusal or a local failure. The capability is remembered per disk. - An all-live page ends the pass, so a quiet pool still costs one LIST per round. - The CAS path honours the storage's batch limit; it used to ignore `objects_chunk_size_to_delete`. Also: `IObjectStorage::batchDeleteKeyLimit`, the S3 adapter keeps the error name on a refused bulk delete, `ThreadName::CAS_GC_JANITOR`, user docs for phases 15-17 and `system.cas_gc_log`. Testing: `CAS*` gate 2605; new integration test `tests/integration/test_cas_janitor_drain` (RustFS; batch, 100-key chunks, no batch delete) fails on the base commit and passes here; A/B on the local RustFS stand (docs/superpowers/reports/2026-10-01-cas-29-9-janitor-ab): 20k keys drain in 1 janitor round instead of 16, 31 s instead of 1307 s, 52 DELETE requests instead of 16k, 160 LIST instead of 3134; on a store without batch delete at 20 ms RTT: 2 rounds instead of 1000 keys per round. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Mikhail Filimonov <mfilimonov@altinity.com>
filimonov
force-pushed
the
fix/antalya-26.6/cas-gc-janitor-batches
branch
from
October 1, 2026 19:47
303a40c to
22afe53
Compare
11 of 30 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The namespace janitor (GC phase 16) deleted a dropped table's dead
_log/_snapkeys one page per folding round, one request per key: a table with 20k parts took 16 rounds and 16kDELETErequests to drain. It now deletes a page's dead keys in batches sized by the store's batch-delete limit, runs the batches as jobs on the GC I/O pool while the next page is listed, and keeps taking pages until a soft 20 s budget.Measured (local RustFS stand, base
e2dd0f5d343vs this branch, 3 runs per cell,DROP TABLEof a table with a long reference stream; full reportdocs/superpowers/reports/2026-10-01-cas-29-9-janitor-ab/README.md):DELETErequestsLISTrequestsDELETE/LISTInside the janitor the page
LISTis now 95–98 % of the per-page time; the batchDELETEis single- to double-digit milliseconds per round. Steady state (no dead keys) stays at oneLISTper round.Changes
_log/_snapkeys per page, sized byIObjectStorage::batchDeleteKeyLimit(objects_chunk_size_to_deleteon S3 — the CAS path used to ignore it — 1 on other native stores, 1000 in the emulated backend); exact-token deletes for_ckpt/_files. Write-once cohorts carry no per-key precondition and no per-key fallback: a batch that fails after retries is leaked and the cursor advances.cas_gc_io_concurrencythreads bound what runs); the next page'sLISToverlaps the running jobs; one-key jobs when the limit is 1.NOT_IMPLEMENTED, remembered per disk), a schedule refusal or a local failure.LISTper round.ThreadName::CAS_GC_JANITOR; docs for phases 15–17,system.cas_gc_log,configuration.md.Risks / notes
DROPis dominated by phases 15 and 17 (manifest and ref deletes, one request per key, serial): ~1074 s for 20k parts at 20 ms RTT, on base and here alike. Follow-up: CAS-337 (same pool for those phases), CAS-338 (~4 GETs per retired part in the fold intake).Testing
CAS*gate 2605 (newCASNamespaceJanitor*suites: batches, failure classification, pipeline window/stop flag/authority latch, wrap, budget; each rule proven by a failing-first test).tests/integration/test_cas_janitor_drainon RustFS — batch disk, 100-key chunks, no batch delete — fails on the base commit (keys left after 3 rounds) and passes here (exactly one deleting round, the one after the fold).DELETEbar is missed (base sends ~200 extra non-janitor deletes in one round, family not identified), every other bar passes.Changelog category (leave one):
Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md):
CAS GC drains a dropped table's reference stream in one round: dead stream keys are deleted in batches sized by the store's limit, on the GC I/O pool, under a soft time budget instead of one page per round.
Documentation entry for user-facing changes
docs/en/antalya/cas/architecture/garbage-collection.md,docs/en/operations/system-tables/cas_gc_log.md,docs/en/antalya/cas/configuration.md)CI/CD Options
Regression jobs to run: