You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 528398b
Browse filesBrowse the repository at this point in the historyBrowse files
CAS GC: batch the namespace janitor's deletes under a soft time budget
The namespace janitor (GC phase 16) deleted dead `_log`/`_snap` keys one
page per folding round: a dropped table with 20k parts took 16 rounds and
16k DELETE requests to drain. It now deletes a page's dead keys in batches
sized by the store's batch-delete limit (`objects_chunk_size_to_delete` on
S3, 1 elsewhere), runs the batches as jobs on the GC I/O pool with at most
`cas_gc_io_concurrency` in flight while the next page is listed, and keeps
taking pages until a soft 20 s budget: no new page after it, in-flight
jobs complete. A pass that began mid-stream wraps once to the stream start.
Rules:
- Dead `_log`/`_snap` keys go in write-once cohorts with no per-key
precondition and no per-key fallback; a batch that still fails after
retries is leaked and the cursor advances.
- A page is held (cursor not advanced, nothing published) only on lost
authority, a capability refusal (`NOT_IMPLEMENTED`), a schedule refusal
or a local failure. The capability is remembered per disk.
- An all-live page ends the pass, so a quiet pool still costs one LIST per
round.
- The CAS path honours the storage's batch limit; it used to ignore
`objects_chunk_size_to_delete`.
Also: `IObjectStorage::batchDeleteKeyLimit`, the S3 adapter keeps the error
name on a refused bulk delete, two ProfileEvents
(`CASGCNamespaceCleanupObjectsDeleted`, `CASGCNamespaceCleanupBatchFailures`),
`ThreadName::CAS_GC_JANITOR`, user docs for phases 15-17 and
`system.cas_gc_log`.
Testing: `CAS*` gate 2628; new integration test
`tests/integration/test_cas_janitor_drain` (RustFS; batch, 100-key chunks,
no batch delete) fails on the base commit and passes here; A/B on the
local RustFS stand (docs/superpowers/reports/2026-10-01-cas-29-9-janitor-ab):
20k keys drain in 1 janitor round instead of 16, 31 s instead of 1307 s,
52 DELETE requests instead of 16k, 160 LIST instead of 3134; on a store
without batch delete at 20 ms RTT: 2 rounds instead of 1000 keys per round.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Mikhail Filimonov <mfilimonov@altinity.com>
Copy file name to clipboardExpand all lines: docs/en/antalya/cas/architecture/garbage-collection.md
+67-41Lines changed: 67 additions & 41 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -57,15 +57,15 @@ follower or a deferred round execution returns before that commit.
57
57
| 13 |`round_commit`| fold | Retention-prune old generations, then publish the single `gc/state``CAS` that adopts the whole round |
58
58
| 14 |`handoff_reclaim`| post-`CAS`| Reclaim a generation a ref moved off during this round, which the ordinary retention prune already skipped and will not revisit |
59
59
| 15 |`manifest_deletes`| post-`CAS`| Delete manifest bodies whose owner-removal minus-one edge the `CAS` in phase 13 just adopted |
60
-
| 16 |`namespace_cleanup`| leader; suppressed on `DEFER`|One bounded page of the perpetual namespace janitor, reclaiming dead-life debris |
60
+
| 16 |`namespace_cleanup`| leader; one suppressed page on `DEFER`|The perpetual namespace janitor: pages of dead-life debris under a 20 s soft budget|
61
61
| 17 |`ref_object_cleanup`| post-`CAS`| Prune ref logs and snapshots once both fold coverage and a live snapshot make them safe to delete |
62
62
| 18 |`orphan_sweep`| post-`CAS`| Exact-token deletion for the [orphan-manifest sweep](/antalya/cas/architecture/manifests-and-refs#orphan-sweep), after phase 13 adopted each candidate's blob-source retirements and the cursor |
63
63
64
64
Phases 2 through 4 run on every leader round; a follower returns after phase 1. Phases 5–15 and
65
65
17–18 run only when phase 4 decides to fold; phase 16 runs after phase 15 on a fold and right after
66
-
phase 4 on a `DEFER`. A `DEFER` verdict is therefore not a bare no-op: it still runs one bounded
67
-
namespace-janitor page with `suppress_destructive = true` — listing and classification only, no
68
-
deletes and no cursor advance — and then returns, publishing no fold artifact and no commit `CAS`.
66
+
phase 4 on a `DEFER`. A folding round runs janitor pages until the phase budget ends; a `DEFER`
67
+
verdict is therefore not a bare no-op: it still runs one namespace-janitor page with `suppress_destructive = true` — it lists and parses the keys but never
68
+
classifies them against the catalog, deletes nothing and does not advance the cursor — and then returns, publishing no fold artifact and no commit `CAS`.
69
69
Its lease `CAS` may already have created or renewed the lease in phase 1:
70
70
71
71
```mermaid
@@ -229,7 +229,7 @@ if any of: changed rows ≥ `gc_fold_threshold` (default 1); an adopted shard ha
229
229
delete; an adopted shard has a condemned blob due to graduate
230
230
(`oldest_nonpending_condemn_round < round + 1`); or `gc_fold_max_defer_rounds` (default 8)
231
231
consecutive defers were reached. Both thresholds are internal `PoolConfig` fields, not disk
232
-
settings. On `defer`, one suppressed namespace-janitor page runs (phase 16's work) and the round
232
+
settings. On `defer`, one suppressed namespace-janitor page runs (phase 16's work, a single page where a folding round runs several) and the round
233
233
returns without a commit. On `fold`, phase 6 reuses this plan and the same `LIST`.
234
234
235
235
## Phase 5 — parent seal read {#phase-5-parent-seal-read}
@@ -500,7 +500,8 @@ Deletes owner-removed manifest bodies, now that phase 13's `CAS` adopted their m
500
500
-**Reads / writes:** batch `DELETE` of the manifest keys collected by phase 8's fold of `-1` owner
501
501
edges, in chunks of `cas_gc_bulk_delete_chunk_keys` (default 1000, the backend maximum). A manifest
502
502
key is write-once, so the delete carries no per-key precondition; an absent key is simply gone. A
503
-
backend without a batch-delete verb (GCS) falls back to one admitted `DELETE` per key.
503
+
storage that rejects a batch delete (`NOT_IMPLEMENTED`) stops the family for the round; the remaining
504
+
bodies are left to the orphan-manifest sweep (phase 18) and later requests carry one key.
504
505
-**Safety:** each body is unreachable from any live ref (its owner-removal was folded and
505
506
committed) and is never re-derived — the intake cursor that found the `-1` edge is now committed,
506
507
so a folded log is never revisited. Hence the phase is unbudgeted by design and drains the whole
@@ -510,34 +511,56 @@ Deletes owner-removed manifest bodies, now that phase 13's `CAS` adopted their m
510
511
all-or-nothing per request: the chunks before the failing one are recorded, the failing chunk's
511
512
keys are not, and a key one of its attempts did delete shows up as already gone in the next fold
512
513
-**Observability:** phase row `manifest_deletes`; metrics `attempted`, `accepted` (keys recorded
513
-
as deleted or absent), `requests` (one per chunk, or the failed bulk call plus one per key on the
514
-
fallback), `suppressed`; one `ManifestDelete` row per key in `system.cas_log`
514
+
as deleted or absent), `requests` (one per chunk), `unsent` (bodies not deleted this round),
515
+
`capability_learned` (1 when the storage rejected a batch delete), `suppressed`; one `ManifestDelete` row per key in `system.cas_log`
515
516
516
-
Only a crash, `suppress_destructive`or a chunk that exhausted its retries leaves an entry — it is
517
-
then picked up by the orphan-manifest sweep (phase 18).
517
+
Only a crash, `suppress_destructive`, a rejected batch delete or a chunk that exhausted its retries
518
+
leaves an entry — it is then picked up by the orphan-manifest sweep (phase 18).
The cursor advances only when the whole page was decided under a held fence and an unambiguous
540
-
catalog; under suppression it lists and classifies but deletes nothing and does not advance.
522
+
The perpetual namespace janitor deletes the physical objects of namespace lives no longer in the catalog
523
+
(dead-life debris), page by page from the durable `janitor_cursor`.
524
+
525
+
-**Runs on:** the fold path, pages until the budget ends; on the deferred path, one page right after
526
+
phase 4 with `suppress_destructive` forced on
527
+
-**Per page (round thread):** a `gc/state` read for authority, one `LIST` page (≤ 1000 keys) of `cas/ns/`,
528
+
one ref-catalog read after the `LIST`, exact-token `DELETE`s of dead `_ckpt` and `_files` objects
529
+
-**Dead `_log` / `_snap` (GC I/O pool):** one batch delete per job of up to `cas_gc_bulk_delete_chunk_keys`
530
+
keys, capped by the storage's batch-delete limit (the disk key `objects_chunk_size_to_delete` on S3, 1 on other native object storages, 1000 in the emulated mode).
531
+
With a limit of 1, jobs are per-key. Jobs run on the GC I/O pool with at most `cas_gc_io_concurrency` in
532
+
flight; the next page is listed while they run
533
+
-**Budget:** 20 s, soft. No page starts after it; the page in progress and the jobs in flight finish
534
+
-**Wrap:** a pass that began mid-stream continues once from the start of the stream after the last page. It
535
+
skips keys past the start cursor, which the pass already handled, and ends at the first page that reaches
536
+
the start cursor
537
+
-**Stops early when:** a page had no dead-life debris, the pass ended, authority was lost, a job held
538
+
its page, a later page's read failed, or a page was left undecided (ambiguous catalog, suppression, or
539
+
an exact delete that lost admission)
540
+
-**Writes:** one `CAS` on the maintenance state at the end, with the cursor after the last complete page
541
+
(empty once the pass has wrapped); nothing under suppression
542
+
-**Failed jobs:** a batch that fails after its retries is leaked: its keys are retried on the next pass
543
+
and the page advances. Lost authority, a refused batch delete (`NOT_IMPLEMENTED`) or a local failure
544
+
holds the page: the cursor stops before it and nothing is published for it. The batch size is fixed
545
+
per phase, so a refusal ends the phase; the capability is remembered per disk and later rounds use
546
+
one-key jobs
547
+
-**Safety:** authority is re-read once per page and each job samples the result before its request, so no job
548
+
sends after the round has observed the loss of authority (a job already past its first request completes);
549
+
every deleted key belongs to a life absent from a catalog cut taken after its page's `LIST`. The incarnation segment in
550
+
every key makes an old life's objects unreachable from a re-created same-name table, so a missed key
551
+
can only leak storage, never expose it
552
+
-**Fails the round if:** nothing; the phase is wrapped in a catch-all ("namespace janitor stopped this
553
+
round")
554
+
-**Observability:** phase row `namespace_cleanup`, see
@@ -753,7 +777,7 @@ folding round is one `LIST` of `cas/ns/stream/`, the heartbeat floor (`LIST` plu
753
777
seal, catalog and `gc/state` reads of phases 2, 4, 5 and 7, one successful lease `CAS`, and one
754
778
commit `CAS`. A deferred round execution is cheaper: the same `LIST`, the heartbeat floor, phase 2's
755
779
seal / catalog / `gc/state` reads, phase 4's two seal reads and catalog read, the lease `GET`/`CAS`,
756
-
and one suppressed namespace-janitor page (its own `LIST` page and reads, no deletes) — no commit
780
+
and one suppressed namespace-janitor page (its own `LIST` page and reads, no deletes; a folding round runs pages until the 20 s budget ends) — no commit
757
781
`CAS` at all.
758
782
759
783
The round's work is self-regulated: what a pass cannot finish within its budgets is carried and
@@ -924,28 +948,30 @@ the hand-off's own budget (`cas_gc_round_handoff_prefix_wholesale_budget`).
924
948
925
949
### Phase 15 — manifest deletes {#cost-phase-15}
926
950
927
-
One batch `DELETE` request per `cas_gc_bulk_delete_chunk_keys` entries of `mf_cleanup` (on a
928
-
backend without batch delete: the refused bulk call plus one `DELETE` per key). No writes under
951
+
One batch `DELETE` request per `min(cas_gc_bulk_delete_chunk_keys, storage limit)` entries of a cohort of
952
+
`mf_cleanup` (a cohort of 1000 keys with a storage limit of 100 is 10 requests; a storage that
953
+
rejects a batch delete ends the family for the round). No writes under
|`<pool_prefix>/cas/ref_catalog`|`GET`|one per page|
938
963
|`<pool_prefix>/gc/state`|`GET`| one per fence check |
939
-
| dead-life object |`DELETE`| one per object (plus one `HEAD` per object whose `LIST` entry carried no token) |
940
-
|`<pool_prefix>/gc/maintenance_state`|`CAS`| 1 when the page is decided |
964
+
| dead `_ckpt` / `_files` object |`DELETE`| one per object (plus one `HEAD` per object whose `LIST` entry carried no token) |
965
+
| dead `_log` / `_snap` keys | batch `DELETE`| one per job of up to `cas_gc_bulk_delete_chunk_keys` keys, capped by the storage's limit; one per key when the limit is 1 |
966
+
|`<pool_prefix>/gc/maintenance_state`|`CAS`| one per phase |
| checkpoint-named `_log`, predecessor seal, `_snap`|`GET`| per planned namespace (recovery-triple validation before any delete) |
947
973
|`<pool_prefix>/cas/ref_catalog` and `<pool_prefix>/gc/state`|`GET`| one each per chunk (authority re-validation) |
948
-
|`_log` / `_snap` keys | batch `DELETE`| one request per chunk of ≤ `cas_gc_bulk_delete_chunk_keys` keys (the refused bulk call plus one per key on a backend without batch delete) |
974
+
|`_log` / `_snap` keys | batch `DELETE`| one request per ≤ `min(cas_gc_bulk_delete_chunk_keys, storage limit)` keys (a chunk of 1000 keys with a storage limit of 100 is 10 requests; a rejected batch delete ends the pass for the round) |
|`cas_gc_meta_pool_size`|`16`| Bounded pool size for GC per-hash freshness-meta writes |
109
-
|`cas_gc_io_concurrency`|`16`| Bounded pool size for GC object-storage requests that run in parallel: the fold's read-ahead (checkpoints, ref logs, manifests, zero-candidate HEADs), the orphan-manifest sweep planning reads, the `SYSTEM CAS GC REBUILD` read-ahead, and the `pending_deletes` blob `HEAD` + conditional `DELETE` fan-out. Not covered: meta writes (`cas_gc_meta_pool_size`) and all other GC requests, which run on the round thread. `1` runs the covered requests sequentially. `cas_gc_read_concurrency` is rejected without an alias; use `cas_gc_io_concurrency` instead |
109
+
|`cas_gc_io_concurrency`|`16`| Bounded pool size for GC object-storage requests that run in parallel: the fold's read-ahead (checkpoints, ref logs, manifests, zero-candidate HEADs), the orphan-manifest sweep planning reads, the `SYSTEM CAS GC REBUILD` read-ahead, the `pending_deletes` blob `HEAD` + conditional `DELETE` fan-out, and the namespace janitor's dead-life deletes. Not covered: meta writes (`cas_gc_meta_pool_size`) and all other GC requests, which run on the round thread. `1` runs the covered requests sequentially. `cas_gc_read_concurrency` is rejected without an alias; use `cas_gc_io_concurrency` instead |
110
110
| `cas_attempt_timeout_ms` | `5000` | Budget for one HTTP attempt of a writable Native mount's control-plane requests (read, head, list, remove, conditional write), at least 1. Together with the connect cap it forms the attempt envelope (`cas_attempt_timeout_ms + 2 × cap`; the cap is `cas_attempt_timeout_ms` itself when the disk's `connect_timeout_ms` is `0`, else `min(connect_timeout_ms, cas_attempt_timeout_ms)`) that the lease arithmetic reserves: one TCP connect and one TLS handshake under the cap each, send/receive bounded per socket operation by `cas_attempt_timeout_ms`. With background renewal the cadence check requires `cas_mount_renew_period_ms + 2 × envelope + cas_lease_safety_margin_ms < cas_mount_lease_ttl_ms`, which puts an effective ceiling on the frozen connect cap: under the defaults (TTL 30000, period 10000, margin 2000) the envelope must stay under 9000, so a disk `connect_timeout_ms` of 2000 ms or more refuses to open writable — lower the connect timeout or raise the TTL if you hit this |
111
111
|`cas_lease_safety_margin_ms`|`2000`| Startup-only margin validated against the mount lease TTL: the attempt envelope + `cas_lease_safety_margin_ms` must be strictly less than the mount lease TTL, and `cas_mount_renew_period_ms` + 2 × envelope + `cas_lease_safety_margin_ms` too, or the disk refuses to open writable |
112
112
|`cas_unsafe_remount_no_delay`|`0`| Reclaim a mount slot that carries this server's own uuid at once after a hard restart, without observing the slot's token for the lease TTL. Unsafe whenever two processes can hold the same `server_uuid` (a copied uuid file, a stalled predecessor). After such a reclaim the predecessor can still start conditional writes until its own cutoff (`confirmed deadline − cas_lease_safety_margin_ms − 2 × envelope`) or until its next renewal meets the token guard, and a request it already sent may still materialize later. That is not a data hazard: ref-log keys carry `(writer_epoch, sequence)` and creates are conditional, so two writers can never commit different bodies to one key, and recovery's epoch seal settles any straggler (recovery fails closed after 64 successive seal-create attempts displaced by newly materializing old-epoch transactions). The exposure is availability, not data. Intended for test stands and deployments that guarantee one process per uuid |
@@ -179,7 +179,7 @@ for the remaining caps, `0` means unbounded.
179
179
|---|---|---|---|
180
180
|`cas_manifest_sweep_list_budget_keys`|`1000`|`UInt64`| Orphan-manifest sweep `LIST` budget per round |
181
181
|`cas_manifest_sweep_delete_budget_keys`|`100`|`UInt64`| Orphan-manifest sweep `DELETE` budget per round |
182
-
|`cas_gc_bulk_delete_chunk_keys`|`1000`|`1`–`1000`| Keys per batch delete request in GC's write-once families (owner-removed manifest bodies, covered ref logs and snapshots)|
182
+
|`cas_gc_bulk_delete_chunk_keys`|`1000`|`1`–`1000`| Keys per batch delete request in GC's write-once families: owner-removed manifest bodies, covered ref logs and snapshots, and dead-life ref logs and snapshots of the namespace janitor. Capped by the storage's batch-delete limit (the disk key `objects_chunk_size_to_delete` on S3, where 0 acts as 1); a storage without batch delete gets one-key requests|
183
183
|`cas_gc_round_graduation_budget`|`5000`|`0` = unbounded | Blob-graduation (`condemned` → `delete_pending`) cohort cap per round |
184
184
|`cas_gc_round_redelete_budget`|`5000`|`0` = unbounded | Exact-token re-delete cohort cap for prior `delete_pending` rows per round |
185
185
|`cas_gc_round_sweep_namespace_budget`|`20`|`0` = unbounded | Distinct namespaces per orphan-manifest sweep page whose protection view may be built |
0 commit comments