Skip to content

CAS GC: batch the namespace janitor's deletes under a soft time budget - #2468

Draft
filimonov wants to merge 1 commit into
antalya-26.6from
fix/antalya-26.6/cas-gc-janitor-batches
Draft

filimonov wants to merge 1 commit into
antalya-26.6from
fix/antalya-26.6/cas-gc-janitor-batches

Conversation

@filimonov

@filimonov filimonov commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

The namespace janitor (GC phase 16) deleted a dropped table's dead _log/_snap keys one page per folding round, one request per key: a table with 20k parts took 16 rounds and 16k DELETE requests to drain. It now deletes a page's dead keys in batches sized by the store's batch-delete limit, runs the batches as jobs on the GC I/O pool while the next page is listed, and keeps taking pages until a soft 20 s budget.

Measured (local RustFS stand, base e2dd0f5d343 vs this branch, 3 runs per cell, DROP TABLE of a table with a long reference stream; full report docs/superpowers/reports/2026-10-01-cas-29-9-janitor-ab/README.md):

Scenario Metric Base This PR
20k keys, 0 ms RTT janitor rounds to drain 16 1
wall time to drain 1307 s 31 s
DELETE requests 16 172–16 194 52 (16 batches)
LIST requests 3 134 160
7k keys, 0 ms rounds / DELETE / LIST 3 / 2 541 / 75 1 / 20 / 44
7k keys, 20 ms RTT janitor phase in the deleting round 22 s 1.2 s
100k keys, 0 ms rounds to drain (this PR only) — 2 (budget stops the first after 20 s; 44–57 pages per round)
20k keys, no batch delete, 20 ms RTT janitor rounds to drain 5 at the cap, 1 000 keys per round 2 (11 000 + 4 959 keys on a 16-thread pool)

Inside the janitor the page LIST is now 95–98 % of the per-page time; the batch DELETE is single- to double-digit milliseconds per round. Steady state (no dead keys) stays at one LIST per round.

Changes

  • Batches of dead _log/_snap keys per page, sized by IObjectStorage::batchDeleteKeyLimit (objects_chunk_size_to_delete on S3 — the CAS path used to ignore it — 1 on other native stores, 1000 in the emulated backend); exact-token deletes for _ckpt/_files. Write-once cohorts carry no per-key precondition and no per-key fallback: a batch that fails after retries is leaked and the cursor advances.
  • Jobs on the GC I/O pool, the same enqueue-and-wait pattern as the pending-deletes fan-out (the pool's cas_gc_io_concurrency threads bound what runs); the next page's LIST overlaps the running jobs; one-key jobs when the limit is 1.
  • A soft 20 s phase budget: no new page after it, in-flight jobs complete. No count budget.
  • A page is held (cursor not advanced, nothing published) only on lost authority — the pass breaks on the first observed loss and never refreshes after it — a capability refusal (NOT_IMPLEMENTED, remembered per disk), a schedule refusal or a local failure.
  • A pass that began mid-stream wraps once to the stream start, so a drop drains in the round after the fold instead of two rounds later. An all-live page ends the pass: a quiet pool still costs one LIST per round.
  • The S3 adapter keeps the error name on a refused bulk delete; ThreadName::CAS_GC_JANITOR; docs for phases 15–17, system.cas_gc_log, configuration.md.

Risks / notes

  • "Drains in the round after the fold" holds when no all-live page sits between the cursor and the debris; otherwise the janitor still advances one page per folding round.
  • On a store without batch delete the round after a DROP is dominated by phases 15 and 17 (manifest and ref deletes, one request per key, serial): ~1074 s for 20k parts at 20 ms RTT, on base and here alike. Follow-up: CAS-337 (same pool for those phases), CAS-338 (~4 GETs per retired part in the fold intake).
  • No on-S3 format change.

Testing

  • Unit: CAS* gate 2605 (new CASNamespaceJanitor* suites: batches, failure classification, pipeline window/stop flag/authority latch, wrap, budget; each rule proven by a failing-first test).
  • Integration: tests/integration/test_cas_janitor_drain on RustFS — batch disk, 100-key chunks, no batch delete — fails on the base commit (keys left after 3 rounds) and passes here (exactly one deleting round, the one after the fold).
  • A/B on the local RustFS stand: the table above; the spec's 1 % wire-DELETE bar is missed (base sends ~200 extra non-janitor deletes in one round, family not identified), every other bar passes.

Changelog category (leave one):

  • Performance Improvement

Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md):

CAS GC drains a dropped table's reference stream in one round: dead stream keys are deleted in batches sized by the store's limit, on the GC I/O pool, under a soft time budget instead of one page per round.

Documentation entry for user-facing changes

  • Documentation written in this PR (docs/en/antalya/cas/architecture/garbage-collection.md, docs/en/operations/system-tables/cas_gc_log.md, docs/en/antalya/cas/configuration.md)

CI/CD Options

Regression jobs to run:

  • CAS (content-addressed storage; Antalya only)

@filimonov filimonov added the CAS label Oct 1, 2026
@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Workflow [PR], commit [22afe53]

@filimonov
filimonov requested a review from k-morozov October 1, 2026 14:10
@filimonov
filimonov marked this pull request as draft October 1, 2026 14:35
@filimonov
filimonov removed the request for review from k-morozov October 1, 2026 14:35
@filimonov
filimonov force-pushed the fix/antalya-26.6/cas-gc-janitor-batches branch 2 times, most recently from c2fef81 to 303a40c Compare October 1, 2026 16:34
The namespace janitor (GC phase 16) deleted dead `_log`/`_snap` keys one
page per folding round: a dropped table with 20k parts took 16 rounds and
16k DELETE requests to drain. It now deletes a page's dead keys in batches
sized by the store's batch-delete limit (`objects_chunk_size_to_delete` on
S3, 1 elsewhere), runs the batches as jobs on the GC I/O pool (the pool's
`cas_gc_io_concurrency` threads bound them, the same pattern as the
pending-deletes fan-out) while the next page is listed, and keeps
taking pages until a soft 20 s budget: no new page after it, in-flight
jobs complete. A pass that began mid-stream wraps once to the stream start.

Rules:
- Dead `_log`/`_snap` keys go in write-once cohorts with no per-key
  precondition and no per-key fallback; a batch that still fails after
  retries is leaked and the cursor advances.
- A page is held (cursor not advanced, nothing published) only on lost
  authority, a capability refusal (`NOT_IMPLEMENTED`), a schedule refusal
  or a local failure. The capability is remembered per disk.
- An all-live page ends the pass, so a quiet pool still costs one LIST per
  round.
- The CAS path honours the storage's batch limit; it used to ignore
  `objects_chunk_size_to_delete`.

Also: `IObjectStorage::batchDeleteKeyLimit`, the S3 adapter keeps the error
name on a refused bulk delete, `ThreadName::CAS_GC_JANITOR`, user docs for
phases 15-17 and `system.cas_gc_log`.

Testing: `CAS*` gate 2605; new integration test
`tests/integration/test_cas_janitor_drain` (RustFS; batch, 100-key chunks,
no batch delete) fails on the base commit and passes here; A/B on the
local RustFS stand (docs/superpowers/reports/2026-10-01-cas-29-9-janitor-ab):
20k keys drain in 1 janitor round instead of 16, 31 s instead of 1307 s,
52 DELETE requests instead of 16k, 160 LIST instead of 3134; on a store
without batch delete at 20 ms RTT: 2 rounds instead of 1000 keys per round.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Mikhail Filimonov <mfilimonov@altinity.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants