Skip to content

Commit 528398b

Browse files
filimonovclaude
andcommitted
CAS GC: batch the namespace janitor's deletes under a soft time budget
The namespace janitor (GC phase 16) deleted dead `_log`/`_snap` keys one page per folding round: a dropped table with 20k parts took 16 rounds and 16k DELETE requests to drain. It now deletes a page's dead keys in batches sized by the store's batch-delete limit (`objects_chunk_size_to_delete` on S3, 1 elsewhere), runs the batches as jobs on the GC I/O pool with at most `cas_gc_io_concurrency` in flight while the next page is listed, and keeps taking pages until a soft 20 s budget: no new page after it, in-flight jobs complete. A pass that began mid-stream wraps once to the stream start. Rules: - Dead `_log`/`_snap` keys go in write-once cohorts with no per-key precondition and no per-key fallback; a batch that still fails after retries is leaked and the cursor advances. - A page is held (cursor not advanced, nothing published) only on lost authority, a capability refusal (`NOT_IMPLEMENTED`), a schedule refusal or a local failure. The capability is remembered per disk. - An all-live page ends the pass, so a quiet pool still costs one LIST per round. - The CAS path honours the storage's batch limit; it used to ignore `objects_chunk_size_to_delete`. Also: `IObjectStorage::batchDeleteKeyLimit`, the S3 adapter keeps the error name on a refused bulk delete, two ProfileEvents (`CASGCNamespaceCleanupObjectsDeleted`, `CASGCNamespaceCleanupBatchFailures`), `ThreadName::CAS_GC_JANITOR`, user docs for phases 15-17 and `system.cas_gc_log`. Testing: `CAS*` gate 2628; new integration test `tests/integration/test_cas_janitor_drain` (RustFS; batch, 100-key chunks, no batch delete) fails on the base commit and passes here; A/B on the local RustFS stand (docs/superpowers/reports/2026-10-01-cas-29-9-janitor-ab): 20k keys drain in 1 janitor round instead of 16, 31 s instead of 1307 s, 52 DELETE requests instead of 16k, 160 LIST instead of 3134; on a store without batch delete at 20 ms RTT: 2 rounds instead of 1000 keys per round. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Mikhail Filimonov <mfilimonov@altinity.com>
1 parent e2dd0f5 commit 528398b

41 files changed

Lines changed: 4123 additions & 723 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎docs/en/antalya/cas/architecture/garbage-collection.md‎

Lines changed: 67 additions & 41 deletions
Original file line numberDiff line numberDiff line change
@@ -57,15 +57,15 @@ follower or a deferred round execution returns before that commit.
5757
| 13 | `round_commit` | fold | Retention-prune old generations, then publish the single `gc/state` `CAS` that adopts the whole round |
5858
| 14 | `handoff_reclaim` | post-`CAS` | Reclaim a generation a ref moved off during this round, which the ordinary retention prune already skipped and will not revisit |
5959
| 15 | `manifest_deletes` | post-`CAS` | Delete manifest bodies whose owner-removal minus-one edge the `CAS` in phase 13 just adopted |
60-
| 16 | `namespace_cleanup` | leader; suppressed on `DEFER` | One bounded page of the perpetual namespace janitor, reclaiming dead-life debris |
60+
| 16 | `namespace_cleanup` | leader; one suppressed page on `DEFER` | The perpetual namespace janitor: pages of dead-life debris under a 20 s soft budget |
6161
| 17 | `ref_object_cleanup` | post-`CAS` | Prune ref logs and snapshots once both fold coverage and a live snapshot make them safe to delete |
6262
| 18 | `orphan_sweep` | post-`CAS` | Exact-token deletion for the [orphan-manifest sweep](/antalya/cas/architecture/manifests-and-refs#orphan-sweep), after phase 13 adopted each candidate's blob-source retirements and the cursor |
6363

6464
Phases 2 through 4 run on every leader round; a follower returns after phase 1. Phases 5–15 and
6565
17–18 run only when phase 4 decides to fold; phase 16 runs after phase 15 on a fold and right after
66-
phase 4 on a `DEFER`. A `DEFER` verdict is therefore not a bare no-op: it still runs one bounded
67-
namespace-janitor page with `suppress_destructive = true` — listing and classification only, no
68-
deletes and no cursor advance — and then returns, publishing no fold artifact and no commit `CAS`.
66+
phase 4 on a `DEFER`. A folding round runs janitor pages until the phase budget ends; a `DEFER`
67+
verdict is therefore not a bare no-op: it still runs one namespace-janitor page with `suppress_destructive = true` — it lists and parses the keys but never
68+
classifies them against the catalog, deletes nothing and does not advance the cursor — and then returns, publishing no fold artifact and no commit `CAS`.
6969
Its lease `CAS` may already have created or renewed the lease in phase 1:
7070

7171
```mermaid
@@ -229,7 +229,7 @@ if any of: changed rows ≥ `gc_fold_threshold` (default 1); an adopted shard ha
229229
delete; an adopted shard has a condemned blob due to graduate
230230
(`oldest_nonpending_condemn_round < round + 1`); or `gc_fold_max_defer_rounds` (default 8)
231231
consecutive defers were reached. Both thresholds are internal `PoolConfig` fields, not disk
232-
settings. On `defer`, one suppressed namespace-janitor page runs (phase 16's work) and the round
232+
settings. On `defer`, one suppressed namespace-janitor page runs (phase 16's work, a single page where a folding round runs several) and the round
233233
returns without a commit. On `fold`, phase 6 reuses this plan and the same `LIST`.
234234

235235
## Phase 5 — parent seal read {#phase-5-parent-seal-read}
@@ -500,7 +500,8 @@ Deletes owner-removed manifest bodies, now that phase 13's `CAS` adopted their m
500500
- **Reads / writes:** batch `DELETE` of the manifest keys collected by phase 8's fold of `-1` owner
501501
edges, in chunks of `cas_gc_bulk_delete_chunk_keys` (default 1000, the backend maximum). A manifest
502502
key is write-once, so the delete carries no per-key precondition; an absent key is simply gone. A
503-
backend without a batch-delete verb (GCS) falls back to one admitted `DELETE` per key.
503+
storage that rejects a batch delete (`NOT_IMPLEMENTED`) stops the family for the round; the remaining
504+
bodies are left to the orphan-manifest sweep (phase 18) and later requests carry one key.
504505
- **Safety:** each body is unreachable from any live ref (its owner-removal was folded and
505506
committed) and is never re-derived — the intake cursor that found the `-1` edge is now committed,
506507
so a folded log is never revisited. Hence the phase is unbudgeted by design and drains the whole
@@ -510,34 +511,56 @@ Deletes owner-removed manifest bodies, now that phase 13's `CAS` adopted their m
510511
all-or-nothing per request: the chunks before the failing one are recorded, the failing chunk's
511512
keys are not, and a key one of its attempts did delete shows up as already gone in the next fold
512513
- **Observability:** phase row `manifest_deletes`; metrics `attempted`, `accepted` (keys recorded
513-
as deleted or absent), `requests` (one per chunk, or the failed bulk call plus one per key on the
514-
fallback), `suppressed`; one `ManifestDelete` row per key in `system.cas_log`
514+
as deleted or absent), `requests` (one per chunk), `unsent` (bodies not deleted this round),
515+
`capability_learned` (1 when the storage rejected a batch delete), `suppressed`; one `ManifestDelete` row per key in `system.cas_log`
515516

516-
Only a crash, `suppress_destructive` or a chunk that exhausted its retries leaves an entry — it is
517-
then picked up by the orphan-manifest sweep (phase 18).
517+
Only a crash, `suppress_destructive`, a rejected batch delete or a chunk that exhausted its retries
518+
leaves an entry — it is then picked up by the orphan-manifest sweep (phase 18).
518519

519520
## Phase 16 — namespace cleanup {#phase-16-namespace-cleanup}
520521

521-
One bounded page of the perpetual namespace janitor: deletes the physical objects of namespace lives
522-
no longer in the catalog (dead-life debris).
523-
524-
- **Runs on:** fold path here; also on the deferred path right after phase 4 with
525-
`suppress_destructive` forced on
526-
- **Reads:** the durable `janitor_cursor`; one `LIST` page (≤ 1000 keys) of `cas/ns/`; a fresh
527-
ref-catalog snapshot; `gc/state` per fence re-check
528-
- **Writes / deletes:** exact-token `DELETE` per dead-life `_log` / `_snap` / `_ckpt` / `_files`
529-
object; one `CAS` on the maintenance state when the page is decided
530-
- **Safety:** each delete is under a GC fence re-check (`lease.owner` / `lease.seq`) before it and
531-
once at the end; the incarnation segment in every key makes an old life's objects structurally
532-
unreachable from a reborn same-name namespace, so a missed key can only leak storage, never expose
533-
it
534-
- **Fails the round if:** nothing — the whole page is wrapped in a catch-all ("namespace janitor
535-
skipped this round")
536-
- **Observability:** phase row `namespace_cleanup`; metrics `janitor_pages`, `janitor_keys`,
537-
`janitor_deleted`, `leaked`
538-
539-
The cursor advances only when the whole page was decided under a held fence and an unambiguous
540-
catalog; under suppression it lists and classifies but deletes nothing and does not advance.
522+
The perpetual namespace janitor deletes the physical objects of namespace lives no longer in the catalog
523+
(dead-life debris), page by page from the durable `janitor_cursor`.
524+
525+
- **Runs on:** the fold path, pages until the budget ends; on the deferred path, one page right after
526+
phase 4 with `suppress_destructive` forced on
527+
- **Per page (round thread):** a `gc/state` read for authority, one `LIST` page (≤ 1000 keys) of `cas/ns/`,
528+
one ref-catalog read after the `LIST`, exact-token `DELETE`s of dead `_ckpt` and `_files` objects
529+
- **Dead `_log` / `_snap` (GC I/O pool):** one batch delete per job of up to `cas_gc_bulk_delete_chunk_keys`
530+
keys, capped by the storage's batch-delete limit (the disk key `objects_chunk_size_to_delete` on S3, 1 on other native object storages, 1000 in the emulated mode).
531+
With a limit of 1, jobs are per-key. Jobs run on the GC I/O pool with at most `cas_gc_io_concurrency` in
532+
flight; the next page is listed while they run
533+
- **Budget:** 20 s, soft. No page starts after it; the page in progress and the jobs in flight finish
534+
- **Wrap:** a pass that began mid-stream continues once from the start of the stream after the last page. It
535+
skips keys past the start cursor, which the pass already handled, and ends at the first page that reaches
536+
the start cursor
537+
- **Stops early when:** a page had no dead-life debris, the pass ended, authority was lost, a job held
538+
its page, a later page's read failed, or a page was left undecided (ambiguous catalog, suppression, or
539+
an exact delete that lost admission)
540+
- **Writes:** one `CAS` on the maintenance state at the end, with the cursor after the last complete page
541+
(empty once the pass has wrapped); nothing under suppression
542+
- **Failed jobs:** a batch that fails after its retries is leaked: its keys are retried on the next pass
543+
and the page advances. Lost authority, a refused batch delete (`NOT_IMPLEMENTED`) or a local failure
544+
holds the page: the cursor stops before it and nothing is published for it. The batch size is fixed
545+
per phase, so a refusal ends the phase; the capability is remembered per disk and later rounds use
546+
one-key jobs
547+
- **Safety:** authority is re-read once per page and each job samples the result before its request, so no job
548+
sends after the round has observed the loss of authority (a job already past its first request completes);
549+
every deleted key belongs to a life absent from a catalog cut taken after its page's `LIST`. The incarnation segment in
550+
every key makes an old life's objects unreachable from a re-created same-name table, so a missed key
551+
can only leak storage, never expose it
552+
- **Fails the round if:** nothing; the phase is wrapped in a catch-all ("namespace janitor stopped this
553+
round")
554+
- **Observability:** phase row `namespace_cleanup`, see
555+
[`system.cas_gc_log`](/operations/system-tables/cas_gc_log#per-phase-rows); events
556+
`CASGCNamespaceCleanupObjectsDeleted`, `CASGCNamespaceCleanupBatchFailures` and
557+
`CASGCNamespaceCleanupLeaks`
558+
559+
### Unversioned buckets {#phase-16-unversioned-buckets}
560+
561+
A batch delete without a token assumes an unversioned bucket, which the mount probe checks. If versioning
562+
is turned on after mount, the batch reports success while noncurrent versions remain: an empty `LIST`
563+
then proves the namespace drained, not that storage was reclaimed.
541564

542565
## Phase 17 — ref object cleanup {#phase-17-ref-object-cleanup}
543566

@@ -551,8 +574,9 @@ from the catalog.
551574
`gc/state` (authority re-validation). No `HEAD`: `_log` / `_snap` keys are write-once, there is
552575
nothing to re-observe
553576
- **Writes / deletes:** batch `DELETE` of the planned `_log` / `_snap` keys in chunks of
554-
`cas_gc_bulk_delete_chunk_keys` (one admitted `DELETE` per key on a backend without batch
555-
delete); the checkpoint-named snapshot is always retained
577+
`cas_gc_bulk_delete_chunk_keys`; a storage that rejects a batch
578+
delete stops the pass for the round, and the same candidates are recomputed next round. The
579+
checkpoint-named snapshot is always retained
556580
- **Safety:** before each chunk, re-validates: ref-catalog token still equals the fold's catalog
557581
cut, same row and life, unchanged GC fence. The first failure stops the whole pass. The
558582
per-round `cas_gc_round_ref_cleanup_budget` cap counts objects and cuts a chunk to what remains;
@@ -563,7 +587,7 @@ from the catalog.
563587
namespace, but a chunk delete that exhausts its retry policy propagates (see
564588
[post-commit failures](#post-commit-failures))
565589
- **Observability:** phase row `ref_object_cleanup`; metrics `namespaces_planned`, `suppressed`,
566-
`trim_enabled`; `ProfileEvent` `CASRefCleanupObjectsDeleted`
590+
`trim_enabled`, `capability_learned`; `ProfileEvent` `CASRefCleanupObjectsDeleted`
567591

568592
## Phase 18 — orphan sweep {#phase-18-orphan-sweep}
569593

@@ -753,7 +777,7 @@ folding round is one `LIST` of `cas/ns/stream/`, the heartbeat floor (`LIST` plu
753777
seal, catalog and `gc/state` reads of phases 2, 4, 5 and 7, one successful lease `CAS`, and one
754778
commit `CAS`. A deferred round execution is cheaper: the same `LIST`, the heartbeat floor, phase 2's
755779
seal / catalog / `gc/state` reads, phase 4's two seal reads and catalog read, the lease `GET`/`CAS`,
756-
and one suppressed namespace-janitor page (its own `LIST` page and reads, no deletes) — no commit
780+
and one suppressed namespace-janitor page (its own `LIST` page and reads, no deletes; a folding round runs pages until the 20 s budget ends) — no commit
757781
`CAS` at all.
758782

759783
The round's work is self-regulated: what a pass cannot finish within its budgets is carried and
@@ -924,28 +948,30 @@ the hand-off's own budget (`cas_gc_round_handoff_prefix_wholesale_budget`).
924948

925949
### Phase 15 — manifest deletes {#cost-phase-15}
926950

927-
One batch `DELETE` request per `cas_gc_bulk_delete_chunk_keys` entries of `mf_cleanup` (on a
928-
backend without batch delete: the refused bulk call plus one `DELETE` per key). No writes under
951+
One batch `DELETE` request per `min(cas_gc_bulk_delete_chunk_keys, storage limit)` entries of a cohort of
952+
`mf_cleanup` (a cohort of 1000 keys with a storage limit of 100 is 10 requests; a storage that
953+
rejects a batch delete ends the family for the round). No writes under
929954
`suppress_destructive`.
930955

931956
### Phase 16 — namespace cleanup {#cost-phase-16}
932957

933958
| Key | Operation | Requests |
934959
|---|---|---:|
935960
| `<pool_prefix>/gc/maintenance_state` | `GET` | 1 (durable `janitor_cursor`) |
936-
| `<pool_prefix>/cas/ns/` | `LIST` | one page |
937-
| `<pool_prefix>/cas/ref_catalog` | `GET` | 1 |
961+
| `<pool_prefix>/cas/ns/` | `LIST` | one per page |
962+
| `<pool_prefix>/cas/ref_catalog` | `GET` | one per page |
938963
| `<pool_prefix>/gc/state` | `GET` | one per fence check |
939-
| dead-life object | `DELETE` | one per object (plus one `HEAD` per object whose `LIST` entry carried no token) |
940-
| `<pool_prefix>/gc/maintenance_state` | `CAS` | 1 when the page is decided |
964+
| dead `_ckpt` / `_files` object | `DELETE` | one per object (plus one `HEAD` per object whose `LIST` entry carried no token) |
965+
| dead `_log` / `_snap` keys | batch `DELETE` | one per job of up to `cas_gc_bulk_delete_chunk_keys` keys, capped by the storage's limit; one per key when the limit is 1 |
966+
| `<pool_prefix>/gc/maintenance_state` | `CAS` | one per phase |
941967

942968
### Phase 17 — ref object cleanup {#cost-phase-17}
943969

944970
| Key | Operation | Requests |
945971
|---|---|---:|
946972
| checkpoint-named `_log`, predecessor seal, `_snap` | `GET` | per planned namespace (recovery-triple validation before any delete) |
947973
| `<pool_prefix>/cas/ref_catalog` and `<pool_prefix>/gc/state` | `GET` | one each per chunk (authority re-validation) |
948-
| `_log` / `_snap` keys | batch `DELETE` | one request per chunk of ≤ `cas_gc_bulk_delete_chunk_keys` keys (the refused bulk call plus one per key on a backend without batch delete) |
974+
| `_log` / `_snap` keys | batch `DELETE` | one request per ≤ `min(cas_gc_bulk_delete_chunk_keys, storage limit)` keys (a chunk of 1000 keys with a storage limit of 100 is 10 requests; a rejected batch delete ends the pass for the round) |
949975

950976
### Phase 18 — orphan sweep {#cost-phase-18}
951977

‎docs/en/antalya/cas/configuration.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -106,7 +106,7 @@ entirely before release. Treat this table as a snapshot of the current build, no
106106
| `cas_part_folder_cache_max_entry_bytes` | 16 MiB | Oversized part-folder views bypass retention above this size |
107107
| `cas_manifest_decode_cache_bytes` | 128 MiB | Manifest decode cache byte budget (`0` disables) |
108108
| `cas_gc_meta_pool_size` | `16` | Bounded pool size for GC per-hash freshness-meta writes |
109-
| `cas_gc_io_concurrency` | `16` | Bounded pool size for GC object-storage requests that run in parallel: the fold's read-ahead (checkpoints, ref logs, manifests, zero-candidate HEADs), the orphan-manifest sweep planning reads, the `SYSTEM CAS GC REBUILD` read-ahead, and the `pending_deletes` blob `HEAD` + conditional `DELETE` fan-out. Not covered: meta writes (`cas_gc_meta_pool_size`) and all other GC requests, which run on the round thread. `1` runs the covered requests sequentially. `cas_gc_read_concurrency` is rejected without an alias; use `cas_gc_io_concurrency` instead |
109+
| `cas_gc_io_concurrency` | `16` | Bounded pool size for GC object-storage requests that run in parallel: the fold's read-ahead (checkpoints, ref logs, manifests, zero-candidate HEADs), the orphan-manifest sweep planning reads, the `SYSTEM CAS GC REBUILD` read-ahead, the `pending_deletes` blob `HEAD` + conditional `DELETE` fan-out, and the namespace janitor's dead-life deletes. Not covered: meta writes (`cas_gc_meta_pool_size`) and all other GC requests, which run on the round thread. `1` runs the covered requests sequentially. `cas_gc_read_concurrency` is rejected without an alias; use `cas_gc_io_concurrency` instead |
110110
| `cas_attempt_timeout_ms` | `5000` | Budget for one HTTP attempt of a writable Native mount's control-plane requests (read, head, list, remove, conditional write), at least 1. Together with the connect cap it forms the attempt envelope (`cas_attempt_timeout_ms + 2 × cap`; the cap is `cas_attempt_timeout_ms` itself when the disk's `connect_timeout_ms` is `0`, else `min(connect_timeout_ms, cas_attempt_timeout_ms)`) that the lease arithmetic reserves: one TCP connect and one TLS handshake under the cap each, send/receive bounded per socket operation by `cas_attempt_timeout_ms`. With background renewal the cadence check requires `cas_mount_renew_period_ms + 2 × envelope + cas_lease_safety_margin_ms < cas_mount_lease_ttl_ms`, which puts an effective ceiling on the frozen connect cap: under the defaults (TTL 30000, period 10000, margin 2000) the envelope must stay under 9000, so a disk `connect_timeout_ms` of 2000 ms or more refuses to open writable — lower the connect timeout or raise the TTL if you hit this |
111111
| `cas_lease_safety_margin_ms` | `2000` | Startup-only margin validated against the mount lease TTL: the attempt envelope + `cas_lease_safety_margin_ms` must be strictly less than the mount lease TTL, and `cas_mount_renew_period_ms` + 2 × envelope + `cas_lease_safety_margin_ms` too, or the disk refuses to open writable |
112112
| `cas_unsafe_remount_no_delay` | `0` | Reclaim a mount slot that carries this server's own uuid at once after a hard restart, without observing the slot's token for the lease TTL. Unsafe whenever two processes can hold the same `server_uuid` (a copied uuid file, a stalled predecessor). After such a reclaim the predecessor can still start conditional writes until its own cutoff (`confirmed deadline − cas_lease_safety_margin_ms − 2 × envelope`) or until its next renewal meets the token guard, and a request it already sent may still materialize later. That is not a data hazard: ref-log keys carry `(writer_epoch, sequence)` and creates are conditional, so two writers can never commit different bodies to one key, and recovery's epoch seal settles any straggler (recovery fails closed after 64 successive seal-create attempts displaced by newly materializing old-epoch transactions). The exposure is availability, not data. Intended for test stands and deployments that guarantee one process per uuid |
@@ -179,7 +179,7 @@ for the remaining caps, `0` means unbounded.
179179
|---|---|---|---|
180180
| `cas_manifest_sweep_list_budget_keys` | `1000` | `UInt64` | Orphan-manifest sweep `LIST` budget per round |
181181
| `cas_manifest_sweep_delete_budget_keys` | `100` | `UInt64` | Orphan-manifest sweep `DELETE` budget per round |
182-
| `cas_gc_bulk_delete_chunk_keys` | `1000` | `1`–`1000` | Keys per batch delete request in GC's write-once families (owner-removed manifest bodies, covered ref logs and snapshots) |
182+
| `cas_gc_bulk_delete_chunk_keys` | `1000` | `1`–`1000` | Keys per batch delete request in GC's write-once families: owner-removed manifest bodies, covered ref logs and snapshots, and dead-life ref logs and snapshots of the namespace janitor. Capped by the storage's batch-delete limit (the disk key `objects_chunk_size_to_delete` on S3, where 0 acts as 1); a storage without batch delete gets one-key requests |
183183
| `cas_gc_round_graduation_budget` | `5000` | `0` = unbounded | Blob-graduation (`condemned` → `delete_pending`) cohort cap per round |
184184
| `cas_gc_round_redelete_budget` | `5000` | `0` = unbounded | Exact-token re-delete cohort cap for prior `delete_pending` rows per round |
185185
| `cas_gc_round_sweep_namespace_budget` | `20` | `0` = unbounded | Distinct namespaces per orphan-manifest sweep page whose protection view may be built |

0 commit comments

Comments
 (0)