You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 08eedde
Browse filesBrowse the repository at this point in the historyBrowse files
Copy file name to clipboardExpand all lines: docs/en/antalya/cas/architecture/garbage-collection.md
+3-3Lines changed: 3 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -324,7 +324,7 @@ exit — a hold, an unusable checkpoint, the probe budget — leaves it unproven
324
324
gate.
325
325
326
326
**Read-ahead.** The checkpoint, walk-position, manifest-edge and (in phase 9) zero-candidate `HEAD`
327
-
reads are hinted ahead onto a bounded pool (`cas_gc_read_concurrency`, default 16; `1`disables) and
327
+
reads are hinted ahead onto a bounded pool (`cas_gc_io_concurrency`, default 16; `1`runs the reads inline) and
328
328
taken by the walk at exactly the sites, and in exactly the order, of the inline reads, so every
329
329
decision, decode, counter and event stays on the round thread and the phase's semantic metrics do
330
330
not depend on the setting. Two things do: a request a worker performed lands on that worker's
@@ -761,7 +761,7 @@ retried by the next round's cursors — with the one exception of phase 14's han
761
761
is one-shot and leaves its remainder to `cas-fsck`. The per-round budgets are ordinary
762
762
`content_addressed` disk settings, documented under
763
763
[advanced GC pacing settings](/antalya/cas/configuration#advanced-gc-pacing-settings) on the
764
-
configuration page (`cas_gc_meta_pool_size` and `cas_gc_read_concurrency` sit in its main
764
+
configuration page (`cas_gc_meta_pool_size` and `cas_gc_io_concurrency` sit in its main
765
765
[disk-settings table](/antalya/cas/configuration#disk-settings)). `0` means unbounded for every
766
766
`cas_gc_round_*` budget; `cas_manifest_sweep_list_budget_keys = 0` disables the sweep,
767
767
`cas_manifest_sweep_delete_budget_keys = 0` lists without nominating, and the two pool sizes and
@@ -781,7 +781,7 @@ the chunk size reject `0`:
781
781
|`cas_gc_round_sweep_recovery_op_budget`| 5000 | committed-tail ref-log reads the sweep's recovery walk may spend (phase 9) |
782
782
|`cas_gc_bulk_delete_chunk_keys`| 1000 | keys per batch `DELETE` request for write-once families (phases 15, 17); `1` to `1000`|
783
783
|`cas_gc_meta_pool_size`| 16 | bounded pool for condemn-marker writes (phase 12) |
784
-
|`cas_gc_read_concurrency`| 16 | bounded pool for the fold's read-ahead of checkpoints, ref logs, manifest bodies and zero-candidate `HEAD`s (phases 8, 9); `1`disables|
784
+
|`cas_gc_io_concurrency`| 16 | bounded pool for the fold's read-ahead of checkpoints, ref logs, manifest bodies and zero-candidate `HEAD`s (phases 8, 9), the orphan-sweep planning reads (phase 9), the rebuild read-ahead and the `pending_deletes``HEAD` + conditional `DELETE` fan-out (phase 11); other GC requests run on the round thread; `1`runs the covered requests sequentially. `cas_gc_read_concurrency` is rejected without an alias; use `cas_gc_io_concurrency` instead|
785
785
786
786
The fold-batching controls `gc_fold_threshold` (default 1), `gc_fold_max_defer_rounds` (default 8)
787
787
and `gc_frontier_probe_budget` (default unbounded) are internal `PoolConfig` fields with no disk
|`cas_gc_meta_pool_size`|`16`| Bounded pool size for GC per-hash freshness-meta writes |
109
-
|`cas_gc_read_concurrency`|`16`| Bounded pool size for the GC fold's read-ahead of checkpoints, ref logs, manifests and zero-candidate HEADs; `1`disables|
109
+
|`cas_gc_io_concurrency`|`16`| Bounded pool size for GC object-storage requests that run in parallel: the fold's read-ahead (checkpoints, ref logs, manifests, zero-candidate HEADs), the orphan-manifest sweep planning reads, the `SYSTEM CAS GC REBUILD` read-ahead, and the `pending_deletes` blob `HEAD` + conditional `DELETE` fan-out. Not covered: meta writes (`cas_gc_meta_pool_size`) and all other GC requests, which run on the round thread. `1`runs the covered requests sequentially. `cas_gc_read_concurrency` is rejected without an alias; use `cas_gc_io_concurrency` instead|
110
110
| `cas_attempt_timeout_ms` | `5000` | Budget for one HTTP attempt of a writable Native mount's control-plane requests (read, head, list, remove, conditional write), at least 1. Together with the connect cap it forms the attempt envelope (`cas_attempt_timeout_ms + 2 × cap`; the cap is `cas_attempt_timeout_ms` itself when the disk's `connect_timeout_ms` is `0`, else `min(connect_timeout_ms, cas_attempt_timeout_ms)`) that the lease arithmetic reserves: one TCP connect and one TLS handshake under the cap each, send/receive bounded per socket operation by `cas_attempt_timeout_ms`. With background renewal the cadence check requires `cas_mount_renew_period_ms + 2 × envelope + cas_lease_safety_margin_ms < cas_mount_lease_ttl_ms`, which puts an effective ceiling on the frozen connect cap: under the defaults (TTL 30000, period 10000, margin 2000) the envelope must stay under 9000, so a disk `connect_timeout_ms` of 2000 ms or more refuses to open writable — lower the connect timeout or raise the TTL if you hit this |
111
111
|`cas_lease_safety_margin_ms`|`2000`| Startup-only margin validated against the mount lease TTL: the attempt envelope + `cas_lease_safety_margin_ms` must be strictly less than the mount lease TTL, and `cas_mount_renew_period_ms` + 2 × envelope + `cas_lease_safety_margin_ms` too, or the disk refuses to open writable |
112
112
|`cas_unsafe_remount_no_delay`|`0`| Reclaim a mount slot that carries this server's own uuid at once after a hard restart, without observing the slot's token for the lease TTL. Unsafe whenever two processes can hold the same `server_uuid` (a copied uuid file, a stalled predecessor). After such a reclaim the predecessor can still start conditional writes until its own cutoff (`confirmed deadline − cas_lease_safety_margin_ms − 2 × envelope`) or until its next renewal meets the token guard, and a request it already sent may still materialize later. That is not a data hazard: ref-log keys carry `(writer_epoch, sequence)` and creates are conditional, so two writers can never commit different bodies to one key, and recovery's epoch seal settles any straggler (recovery fails closed after 64 successive seal-create attempts displaced by newly materializing old-epoch transactions). The exposure is availability, not data. Intended for test stands and deployments that guarantee one process per uuid |
| 1 | server_name | String | universal | always | Server identifier. An Altinity Antalya build appends `" (antalya:N)"`, where `N` is its Antalya protocol version; a client may ignore or strip the suffix. See [Antalya protocol version](/antalya/protocol).|
444
444
| 2 | version_major | VarUInt | universal | always | Server major version |
445
445
| 3 | version_minor | VarUInt | universal | always | Server minor version |
-`entries_graduated` ([UInt64](/sql-reference/data-types/int-uint)) — Retired entries newly floor-passed and republished `delete_pending` this round (pipeline stage 2; deleted the next round).
-`entries_redelete_failed` ([UInt64](/sql-reference/data-types/int-uint)) — Pending blob deletes whose `HEAD` or exact-token `DELETE` failed this round. Each such entry stays `delete_pending` and is retried in the next round; a non-zero value fails the round.
51
52
-`fence_outs` ([UInt64](/sql-reference/data-types/int-uint)) — Expired mounts fenced out by this round's heartbeat floor.
52
53
-`anomalies` ([UInt64](/sql-reference/data-types/int-uint)) — Fold clamps surfaced (and survived) this round. A steady non-zero value warrants a look at the round log details.
53
54
-`duration_ms` ([UInt64](/sql-reference/data-types/int-uint)) — The round wall-clock duration (on a `Finish` row).
@@ -80,7 +81,7 @@ The phases, in execution order:
80
81
|`fold_ref_intake`| Read and fold every new ref log and the manifest bodies its edges name. | one `GET` per new log, one `GET` per manifest edge |
81
82
|`fold_reduce`| The per-shard in-degree merge: condemn, spare, graduate. | prior-run streaming `GET`s, one `HEAD` per zero-transition candidate, run `PUT`s |
82
83
|`fold_seal_write`| Publish the new fold seal. | one `PUT`|
83
-
|`pending_deletes`| The single content-delete site: exact-token deletes of previously published `delete_pending` entries, plus the outcome logs. | one `DELETE` per entry, one outcome-log `PUT` per shard |
84
+
|`pending_deletes`| The single content-delete site: exact-token deletes of previously published `delete_pending` entries, plus the outcome logs. `phase_metrics` carries `jobs_scheduled` (entries sent to the GC I/O pool; `0` for no candidates, sequential execution, singleton batches, or refusal before the first submission) and `jobs_failed`. | one `HEAD` and at most one `DELETE` per entry, one outcome-log `PUT` per shard with rows|
84
85
|`meta_pool_wait`| Drain the round's per-hash freshness-meta writes. | none on this thread — see the caveat below |
85
86
|`round_commit`| The generation-retention prune and the round's single `gc/state` compare-and-swap. | prune `LIST`s and deletes, one compare-and-swap |
86
87
|`handoff_reclaim`| Wholesale-reclaim generations a moved run ref stranded below the retention cursor. | prefix `LIST`s and deletes |
@@ -128,6 +129,10 @@ Two caveats when reading these rows:
128
129
- Work scheduled onto the GC meta pool runs on other threads, so the `meta_pool_wait` row's
129
130
`ProfileEvents` delta is **empty by construction**. Read its `phase_metrics``jobs_scheduled` /
130
131
`jobs_completed` next to its duration instead: they distinguish a deep queue from a slow endpoint.
132
+
- Requests issued by GC I/O pool workers — the `pending_deletes` fan-out and the fold read-ahead of
133
+
`fold_ref_intake` and `fold_reduce` — also run on other threads and are missing from those phase
134
+
rows' `ProfileEvents`. They still count in `system.events`. On `pending_deletes`, read
135
+
`phase_metrics``jobs_scheduled` / `jobs_failed` next to `redeleted`.
131
136
- Phase durations do not sum to the round's `duration_ms`. The round also performs untimed
132
137
bookkeeping between phases, and the `Finish` row's `duration_ms` remains the authority on total
Copy file name to clipboardExpand all lines: docs/en/sql-reference/statements/system.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -479,7 +479,7 @@ When `disk_name` is given, the round runs on that content-addressed disk only; t
479
479
480
480
Each round is recorded in [`system.cas_gc_log`](/operations/system-tables/cas_gc_log) as a `Start`and a `Finish` row (with `trigger = 'Manual'`).
481
481
482
-
The command returns one row per disk it ran on (multiple rows when `disk_name` is omitted), with columns `disk`, `acquired_lease`, `deferred`, `round`, `candidates_marked`, `objects_deleted`, `objects_absent`, `objects_replaced`, `objects_spared`, `manifests_deleted`, `entries_condemned`, `entries_graduated`, `entries_redeleted`, `fence_outs`, `anomalies`, `pending_candidates`, `pending_condemned`, and`pending_retired`, describing the outcome of that round. The `pending_*` columns are the retire pipeline's remaining backlog sizes read from the `gc/state` this round's own commit just published (not this round's own delta, unlike the columns before them) — `0` on a non-authoritative row (`acquired_lease = 0` or `deferred = 1`), same as every other counter.
482
+
The command returns one row per disk it ran on (multiple rows when `disk_name` is omitted), with columns `disk`, `acquired_lease`, `deferred`, `round`, `candidates_marked`, `objects_deleted`, `objects_absent`, `objects_replaced`, `objects_spared`, `manifests_deleted`, `entries_condemned`, `entries_graduated`, `entries_redeleted`, `entries_redelete_failed`, `fence_outs`, `anomalies`, `pending_candidates`, `pending_condemned`, and`pending_retired`, describing the outcome of that round. The `pending_*` columns are the retire pipeline's remaining backlog sizes read from the `gc/state` this round's own commit just published (not this round's own delta, unlike the columns before them) — `0` on a non-authoritative row (`acquired_lease = 0` or `deferred = 1`), same as every other counter.
483
483
484
484
A manual run always executes, regardless of [`SYSTEM CAS GC STOP`](#system-cas-gc-stop-start): `STOP` pauses only the background scheduler on that disk.
0 commit comments