Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
108 changes: 67 additions & 41 deletions docs/en/antalya/cas/architecture/garbage-collection.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,15 +57,15 @@ follower or a deferred round execution returns before that commit.
| 13 | `round_commit` | fold | Retention-prune old generations, then publish the single `gc/state` `CAS` that adopts the whole round |
| 14 | `handoff_reclaim` | post-`CAS` | Reclaim a generation a ref moved off during this round, which the ordinary retention prune already skipped and will not revisit |
| 15 | `manifest_deletes` | post-`CAS` | Delete manifest bodies whose owner-removal minus-one edge the `CAS` in phase 13 just adopted |
| 16 | `namespace_cleanup` | leader; suppressed on `DEFER` | One bounded page of the perpetual namespace janitor, reclaiming dead-life debris |
| 16 | `namespace_cleanup` | leader; one suppressed page on `DEFER` | The perpetual namespace janitor: pages of dead-life debris under a 20 s soft budget |
| 17 | `ref_object_cleanup` | post-`CAS` | Prune ref logs and snapshots once both fold coverage and a live snapshot make them safe to delete |
| 18 | `orphan_sweep` | post-`CAS` | Exact-token deletion for the [orphan-manifest sweep](/antalya/cas/architecture/manifests-and-refs#orphan-sweep), after phase 13 adopted each candidate's blob-source retirements and the cursor |

Phases 2 through 4 run on every leader round; a follower returns after phase 1. Phases 5–15 and
17–18 run only when phase 4 decides to fold; phase 16 runs after phase 15 on a fold and right after
phase 4 on a `DEFER`. A `DEFER` verdict is therefore not a bare no-op: it still runs one bounded
namespace-janitor page with `suppress_destructive = true` — listing and classification only, no
deletes and no cursor advance — and then returns, publishing no fold artifact and no commit `CAS`.
phase 4 on a `DEFER`. A folding round runs janitor pages until the phase budget ends; a `DEFER`
verdict is therefore not a bare no-op: it still runs one namespace-janitor page with `suppress_destructive = true` — it lists and parses the keys but never
classifies them against the catalog, deletes nothing and does not advance the cursor — and then returns, publishing no fold artifact and no commit `CAS`.
Its lease `CAS` may already have created or renewed the lease in phase 1:

```mermaid
Expand Down Expand Up @@ -229,7 +229,7 @@ if any of: changed rows ≥ `gc_fold_threshold` (default 1); an adopted shard ha
delete; an adopted shard has a condemned blob due to graduate
(`oldest_nonpending_condemn_round < round + 1`); or `gc_fold_max_defer_rounds` (default 8)
consecutive defers were reached. Both thresholds are internal `PoolConfig` fields, not disk
settings. On `defer`, one suppressed namespace-janitor page runs (phase 16's work) and the round
settings. On `defer`, one suppressed namespace-janitor page runs (phase 16's work, a single page where a folding round runs several) and the round
returns without a commit. On `fold`, phase 6 reuses this plan and the same `LIST`.

## Phase 5 — parent seal read {#phase-5-parent-seal-read}
Expand Down Expand Up @@ -500,7 +500,8 @@ Deletes owner-removed manifest bodies, now that phase 13's `CAS` adopted their m
- **Reads / writes:** batch `DELETE` of the manifest keys collected by phase 8's fold of `-1` owner
edges, in chunks of `cas_gc_bulk_delete_chunk_keys` (default 1000, the backend maximum). A manifest
key is write-once, so the delete carries no per-key precondition; an absent key is simply gone. A
backend without a batch-delete verb (GCS) falls back to one admitted `DELETE` per key.
storage that rejects a batch delete (`NOT_IMPLEMENTED`) stops the family for the round; the remaining
bodies are left to the orphan-manifest sweep (phase 18) and later requests carry one key.
- **Safety:** each body is unreachable from any live ref (its owner-removal was folded and
committed) and is never re-derived — the intake cursor that found the `-1` edge is now committed,
so a folded log is never revisited. Hence the phase is unbudgeted by design and drains the whole
Expand All @@ -510,34 +511,56 @@ Deletes owner-removed manifest bodies, now that phase 13's `CAS` adopted their m
all-or-nothing per request: the chunks before the failing one are recorded, the failing chunk's
keys are not, and a key one of its attempts did delete shows up as already gone in the next fold
- **Observability:** phase row `manifest_deletes`; metrics `attempted`, `accepted` (keys recorded
as deleted or absent), `requests` (one per chunk, or the failed bulk call plus one per key on the
fallback), `suppressed`; one `ManifestDelete` row per key in `system.cas_log`
as deleted or absent), `requests` (one per chunk), `unsent` (bodies not deleted this round),
`capability_learned` (1 when the storage rejected a batch delete), `suppressed`; one `ManifestDelete` row per key in `system.cas_log`

Only a crash, `suppress_destructive` or a chunk that exhausted its retries leaves an entry — it is
then picked up by the orphan-manifest sweep (phase 18).
Only a crash, `suppress_destructive`, a rejected batch delete or a chunk that exhausted its retries
leaves an entry — it is then picked up by the orphan-manifest sweep (phase 18).

## Phase 16 — namespace cleanup {#phase-16-namespace-cleanup}

One bounded page of the perpetual namespace janitor: deletes the physical objects of namespace lives
no longer in the catalog (dead-life debris).

- **Runs on:** fold path here; also on the deferred path right after phase 4 with
`suppress_destructive` forced on
- **Reads:** the durable `janitor_cursor`; one `LIST` page (≤ 1000 keys) of `cas/ns/`; a fresh
ref-catalog snapshot; `gc/state` per fence re-check
- **Writes / deletes:** exact-token `DELETE` per dead-life `_log` / `_snap` / `_ckpt` / `_files`
object; one `CAS` on the maintenance state when the page is decided
- **Safety:** each delete is under a GC fence re-check (`lease.owner` / `lease.seq`) before it and
once at the end; the incarnation segment in every key makes an old life's objects structurally
unreachable from a reborn same-name namespace, so a missed key can only leak storage, never expose
it
- **Fails the round if:** nothing — the whole page is wrapped in a catch-all ("namespace janitor
skipped this round")
- **Observability:** phase row `namespace_cleanup`; metrics `janitor_pages`, `janitor_keys`,
`janitor_deleted`, `leaked`

The cursor advances only when the whole page was decided under a held fence and an unambiguous
catalog; under suppression it lists and classifies but deletes nothing and does not advance.
The perpetual namespace janitor deletes the physical objects of namespace lives no longer in the catalog
(dead-life debris), page by page from the durable `janitor_cursor`.

- **Runs on:** the fold path, pages until the budget ends; on the deferred path, one page right after
phase 4 with `suppress_destructive` forced on
- **Per page (round thread):** a `gc/state` read for authority, one `LIST` page (≤ 1000 keys) of `cas/ns/`,
one ref-catalog read after the `LIST`, exact-token `DELETE`s of dead `_ckpt` and `_files` objects
- **Dead `_log` / `_snap` (GC I/O pool):** one batch delete per job of up to `cas_gc_bulk_delete_chunk_keys`
keys, capped by the storage's batch-delete limit (the disk key `objects_chunk_size_to_delete` on S3, 1 on other native object storages, 1000 in the emulated mode).
With a limit of 1, jobs are per-key. A page's jobs are all enqueued on the GC I/O pool, whose
`cas_gc_io_concurrency` threads bound how many run at once; the next page is listed while they run, and
they finish before that page is used
- **Budget:** 20 s, soft. No page starts after it; the page in progress and the jobs in flight finish
- **Wrap:** a pass that began mid-stream continues once from the start of the stream after the last page. It
skips keys past the start cursor, which the pass already handled, and ends at the first page that reaches
the start cursor
- **Stops early when:** a page had no dead-life debris, the pass ended, authority was lost, a job held
its page, a later page's read failed, or a page was left undecided (ambiguous catalog, suppression, or
an exact delete that lost admission)
- **Writes:** one `CAS` on the maintenance state at the end, with the cursor after the last complete page
(empty once the pass has wrapped); nothing under suppression
- **Failed jobs:** a batch that fails after its retries is leaked: its keys are retried on the next pass
and the page advances. Lost authority, a refused batch delete (`NOT_IMPLEMENTED`) or a local failure
holds the page: the cursor stops before it and nothing is published for it. The batch size is fixed
per phase, so a refusal ends the phase; the capability is remembered per disk and later rounds use
one-key jobs
- **Safety:** authority is re-read once per page and each job samples the result before its request, so no job
sends after the round has observed the loss of authority (a job already past its first request completes);
every deleted key belongs to a life absent from a catalog cut taken after its page's `LIST`. The incarnation segment in
every key makes an old life's objects unreachable from a re-created same-name table, so a missed key
can only leak storage, never expose it
- **Fails the round if:** nothing; the phase is wrapped in a catch-all ("namespace janitor stopped this
round")
- **Observability:** phase row `namespace_cleanup`, see
[`system.cas_gc_log`](/operations/system-tables/cas_gc_log#per-phase-rows); event
`CASGCNamespaceCleanupLeaks`

### Unversioned buckets {#phase-16-unversioned-buckets}

A batch delete without a token assumes an unversioned bucket, which the mount probe checks. If versioning
is turned on after mount, the batch reports success while noncurrent versions remain: an empty `LIST`
then proves the namespace drained, not that storage was reclaimed.

## Phase 17 — ref object cleanup {#phase-17-ref-object-cleanup}

Expand All @@ -551,8 +574,9 @@ from the catalog.
`gc/state` (authority re-validation). No `HEAD`: `_log` / `_snap` keys are write-once, there is
nothing to re-observe
- **Writes / deletes:** batch `DELETE` of the planned `_log` / `_snap` keys in chunks of
`cas_gc_bulk_delete_chunk_keys` (one admitted `DELETE` per key on a backend without batch
delete); the checkpoint-named snapshot is always retained
`cas_gc_bulk_delete_chunk_keys`; a storage that rejects a batch
delete stops the pass for the round, and the same candidates are recomputed next round. The
checkpoint-named snapshot is always retained
- **Safety:** before each chunk, re-validates: ref-catalog token still equals the fold's catalog
cut, same row and life, unchanged GC fence. The first failure stops the whole pass. The
per-round `cas_gc_round_ref_cleanup_budget` cap counts objects and cuts a chunk to what remains;
Expand All @@ -563,7 +587,7 @@ from the catalog.
namespace, but a chunk delete that exhausts its retry policy propagates (see
[post-commit failures](#post-commit-failures))
- **Observability:** phase row `ref_object_cleanup`; metrics `namespaces_planned`, `suppressed`,
`trim_enabled`; `ProfileEvent` `CASRefCleanupObjectsDeleted`
`trim_enabled`, `capability_learned`; `ProfileEvent` `CASRefCleanupObjectsDeleted`

## Phase 18 — orphan sweep {#phase-18-orphan-sweep}

Expand Down Expand Up @@ -753,7 +777,7 @@ folding round is one `LIST` of `cas/ns/stream/`, the heartbeat floor (`LIST` plu
seal, catalog and `gc/state` reads of phases 2, 4, 5 and 7, one successful lease `CAS`, and one
commit `CAS`. A deferred round execution is cheaper: the same `LIST`, the heartbeat floor, phase 2's
seal / catalog / `gc/state` reads, phase 4's two seal reads and catalog read, the lease `GET`/`CAS`,
and one suppressed namespace-janitor page (its own `LIST` page and reads, no deletes) — no commit
and one suppressed namespace-janitor page (its own `LIST` page and reads, no deletes; a folding round runs pages until the 20 s budget ends) — no commit
`CAS` at all.

The round's work is self-regulated: what a pass cannot finish within its budgets is carried and
Expand Down Expand Up @@ -924,28 +948,30 @@ the hand-off's own budget (`cas_gc_round_handoff_prefix_wholesale_budget`).

### Phase 15 — manifest deletes {#cost-phase-15}

One batch `DELETE` request per `cas_gc_bulk_delete_chunk_keys` entries of `mf_cleanup` (on a
backend without batch delete: the refused bulk call plus one `DELETE` per key). No writes under
One batch `DELETE` request per `min(cas_gc_bulk_delete_chunk_keys, storage limit)` entries of a cohort of
`mf_cleanup` (a cohort of 1000 keys with a storage limit of 100 is 10 requests; a storage that
rejects a batch delete ends the family for the round). No writes under
`suppress_destructive`.

### Phase 16 — namespace cleanup {#cost-phase-16}

| Key | Operation | Requests |
|---|---|---:|
| `<pool_prefix>/gc/maintenance_state` | `GET` | 1 (durable `janitor_cursor`) |
| `<pool_prefix>/cas/ns/` | `LIST` | one page |
| `<pool_prefix>/cas/ref_catalog` | `GET` | 1 |
| `<pool_prefix>/cas/ns/` | `LIST` | one per page |
| `<pool_prefix>/cas/ref_catalog` | `GET` | one per page |
| `<pool_prefix>/gc/state` | `GET` | one per fence check |
| dead-life object | `DELETE` | one per object (plus one `HEAD` per object whose `LIST` entry carried no token) |
| `<pool_prefix>/gc/maintenance_state` | `CAS` | 1 when the page is decided |
| dead `_ckpt` / `_files` object | `DELETE` | one per object (plus one `HEAD` per object whose `LIST` entry carried no token) |
| dead `_log` / `_snap` keys | batch `DELETE` | one per job of up to `cas_gc_bulk_delete_chunk_keys` keys, capped by the storage's limit; one per key when the limit is 1 |
| `<pool_prefix>/gc/maintenance_state` | `CAS` | one per phase |

### Phase 17 — ref object cleanup {#cost-phase-17}

| Key | Operation | Requests |
|---|---|---:|
| checkpoint-named `_log`, predecessor seal, `_snap` | `GET` | per planned namespace (recovery-triple validation before any delete) |
| `<pool_prefix>/cas/ref_catalog` and `<pool_prefix>/gc/state` | `GET` | one each per chunk (authority re-validation) |
| `_log` / `_snap` keys | batch `DELETE` | one request per chunk of ≤ `cas_gc_bulk_delete_chunk_keys` keys (the refused bulk call plus one per key on a backend without batch delete) |
| `_log` / `_snap` keys | batch `DELETE` | one request per ≤ `min(cas_gc_bulk_delete_chunk_keys, storage limit)` keys (a chunk of 1000 keys with a storage limit of 100 is 10 requests; a rejected batch delete ends the pass for the round) |

### Phase 18 — orphan sweep {#cost-phase-18}

Expand Down
4 changes: 2 additions & 2 deletions docs/en/antalya/cas/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -106,7 +106,7 @@ entirely before release. Treat this table as a snapshot of the current build, no
| `cas_part_folder_cache_max_entry_bytes` | 16 MiB | Oversized part-folder views bypass retention above this size |
| `cas_manifest_decode_cache_bytes` | 128 MiB | Manifest decode cache byte budget (`0` disables) |
| `cas_gc_meta_pool_size` | `16` | Bounded pool size for GC per-hash freshness-meta writes |
| `cas_gc_io_concurrency` | `16` | Bounded pool size for GC object-storage requests that run in parallel: the fold's read-ahead (checkpoints, ref logs, manifests, zero-candidate HEADs), the orphan-manifest sweep planning reads, the `SYSTEM CAS GC REBUILD` read-ahead, and the `pending_deletes` blob `HEAD` + conditional `DELETE` fan-out. Not covered: meta writes (`cas_gc_meta_pool_size`) and all other GC requests, which run on the round thread. `1` runs the covered requests sequentially. `cas_gc_read_concurrency` is rejected without an alias; use `cas_gc_io_concurrency` instead |
| `cas_gc_io_concurrency` | `16` | Bounded pool size for GC object-storage requests that run in parallel: the fold's read-ahead (checkpoints, ref logs, manifests, zero-candidate HEADs), the orphan-manifest sweep planning reads, the `SYSTEM CAS GC REBUILD` read-ahead, the `pending_deletes` blob `HEAD` + conditional `DELETE` fan-out, and the namespace janitor's dead-life deletes. Not covered: meta writes (`cas_gc_meta_pool_size`) and all other GC requests, which run on the round thread. `1` runs the covered requests on a one-thread pool, one at a time; the janitor's jobs still overlap the round thread's authority refresh and next page `LIST`. `cas_gc_read_concurrency` is rejected without an alias; use `cas_gc_io_concurrency` instead |
| `cas_attempt_timeout_ms` | `5000` | Budget for one HTTP attempt of a writable Native mount's control-plane requests (read, head, list, remove, conditional write), at least 1. Together with the connect cap it forms the attempt envelope (`cas_attempt_timeout_ms + 2 × cap`; the cap is `cas_attempt_timeout_ms` itself when the disk's `connect_timeout_ms` is `0`, else `min(connect_timeout_ms, cas_attempt_timeout_ms)`) that the lease arithmetic reserves: one TCP connect and one TLS handshake under the cap each, send/receive bounded per socket operation by `cas_attempt_timeout_ms`. With background renewal the cadence check requires `cas_mount_renew_period_ms + 2 × envelope + cas_lease_safety_margin_ms < cas_mount_lease_ttl_ms`, which puts an effective ceiling on the frozen connect cap: under the defaults (TTL 30000, period 10000, margin 2000) the envelope must stay under 9000, so a disk `connect_timeout_ms` of 2000 ms or more refuses to open writable — lower the connect timeout or raise the TTL if you hit this |
| `cas_lease_safety_margin_ms` | `2000` | Startup-only margin validated against the mount lease TTL: the attempt envelope + `cas_lease_safety_margin_ms` must be strictly less than the mount lease TTL, and `cas_mount_renew_period_ms` + 2 × envelope + `cas_lease_safety_margin_ms` too, or the disk refuses to open writable |
| `cas_unsafe_remount_no_delay` | `0` | Reclaim a mount slot that carries this server's own uuid at once after a hard restart, without observing the slot's token for the lease TTL. Unsafe whenever two processes can hold the same `server_uuid` (a copied uuid file, a stalled predecessor). After such a reclaim the predecessor can still start conditional writes until its own cutoff (`confirmed deadline − cas_lease_safety_margin_ms − 2 × envelope`) or until its next renewal meets the token guard, and a request it already sent may still materialize later. That is not a data hazard: ref-log keys carry `(writer_epoch, sequence)` and creates are conditional, so two writers can never commit different bodies to one key, and recovery's epoch seal settles any straggler (recovery fails closed after 64 successive seal-create attempts displaced by newly materializing old-epoch transactions). The exposure is availability, not data. Intended for test stands and deployments that guarantee one process per uuid |
Expand Down Expand Up @@ -179,7 +179,7 @@ for the remaining caps, `0` means unbounded.
|---|---|---|---|
| `cas_manifest_sweep_list_budget_keys` | `1000` | `UInt64` | Orphan-manifest sweep `LIST` budget per round |
| `cas_manifest_sweep_delete_budget_keys` | `100` | `UInt64` | Orphan-manifest sweep `DELETE` budget per round |
| `cas_gc_bulk_delete_chunk_keys` | `1000` | `1`–`1000` | Keys per batch delete request in GC's write-once families (owner-removed manifest bodies, covered ref logs and snapshots) |
| `cas_gc_bulk_delete_chunk_keys` | `1000` | `1`–`1000` | Keys per batch delete request in GC's write-once families: owner-removed manifest bodies, covered ref logs and snapshots, and dead-life ref logs and snapshots of the namespace janitor. Capped by the storage's batch-delete limit (the disk key `objects_chunk_size_to_delete` on S3, where 0 acts as 1); a storage without batch delete gets one-key requests |
| `cas_gc_round_graduation_budget` | `5000` | `0` = unbounded | Blob-graduation (`condemned` → `delete_pending`) cohort cap per round |
| `cas_gc_round_redelete_budget` | `5000` | `0` = unbounded | Exact-token re-delete cohort cap for prior `delete_pending` rows per round |
| `cas_gc_round_sweep_namespace_budget` | `20` | `0` = unbounded | Distinct namespaces per orphan-manifest sweep page whose protection view may be built |
Expand Down
Loading
Loading