Skip to content

Commit 08eedde

Browse files
authored
Merge branch 'antalya-26.6' into feature/antalya-26.6/aggregate-function-states-in-parquet-iceberg
2 parents 632a76d + e2dd0f5 commit 08eedde

65 files changed

Lines changed: 3475 additions & 155 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎ci/docker/integration/runner/Dockerfile‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -77,7 +77,7 @@ RUN curl -fsSL -O https://archive.apache.org/dist/spark/spark-3.5.5/spark-3.5.5-
7777
# if you change packages, don't forget to update them in tests/integration/helpers/cluster.py
7878
RUN packages="io.delta:delta-spark_2.12:3.1.0,\
7979
org.apache.hudi:hudi-spark3.5-bundle_2.12:1.0.1,\
80-
org.apache.iceberg:iceberg-spark-runtime-3.5_2.12:1.8.1,\
80+
org.apache.iceberg:iceberg-spark-runtime-3.5_2.12:1.10.0,\
8181
org.apache.hadoop:hadoop-aws:3.3.4,\
8282
com.amazonaws:aws-java-sdk-bundle:1.12.262,\
8383
org.apache.hadoop:hadoop-azure:3.3.4,\

‎docs/en/antalya/cas/architecture/garbage-collection.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -324,7 +324,7 @@ exit — a hold, an unusable checkpoint, the probe budget — leaves it unproven
324324
gate.
325325

326326
**Read-ahead.** The checkpoint, walk-position, manifest-edge and (in phase 9) zero-candidate `HEAD`
327-
reads are hinted ahead onto a bounded pool (`cas_gc_read_concurrency`, default 16; `1` disables) and
327+
reads are hinted ahead onto a bounded pool (`cas_gc_io_concurrency`, default 16; `1` runs the reads inline) and
328328
taken by the walk at exactly the sites, and in exactly the order, of the inline reads, so every
329329
decision, decode, counter and event stays on the round thread and the phase's semantic metrics do
330330
not depend on the setting. Two things do: a request a worker performed lands on that worker's
@@ -761,7 +761,7 @@ retried by the next round's cursors — with the one exception of phase 14's han
761761
is one-shot and leaves its remainder to `cas-fsck`. The per-round budgets are ordinary
762762
`content_addressed` disk settings, documented under
763763
[advanced GC pacing settings](/antalya/cas/configuration#advanced-gc-pacing-settings) on the
764-
configuration page (`cas_gc_meta_pool_size` and `cas_gc_read_concurrency` sit in its main
764+
configuration page (`cas_gc_meta_pool_size` and `cas_gc_io_concurrency` sit in its main
765765
[disk-settings table](/antalya/cas/configuration#disk-settings)). `0` means unbounded for every
766766
`cas_gc_round_*` budget; `cas_manifest_sweep_list_budget_keys = 0` disables the sweep,
767767
`cas_manifest_sweep_delete_budget_keys = 0` lists without nominating, and the two pool sizes and
@@ -781,7 +781,7 @@ the chunk size reject `0`:
781781
| `cas_gc_round_sweep_recovery_op_budget` | 5000 | committed-tail ref-log reads the sweep's recovery walk may spend (phase 9) |
782782
| `cas_gc_bulk_delete_chunk_keys` | 1000 | keys per batch `DELETE` request for write-once families (phases 15, 17); `1` to `1000` |
783783
| `cas_gc_meta_pool_size` | 16 | bounded pool for condemn-marker writes (phase 12) |
784-
| `cas_gc_read_concurrency` | 16 | bounded pool for the fold's read-ahead of checkpoints, ref logs, manifest bodies and zero-candidate `HEAD`s (phases 8, 9); `1` disables |
784+
| `cas_gc_io_concurrency` | 16 | bounded pool for the fold's read-ahead of checkpoints, ref logs, manifest bodies and zero-candidate `HEAD`s (phases 8, 9), the orphan-sweep planning reads (phase 9), the rebuild read-ahead and the `pending_deletes` `HEAD` + conditional `DELETE` fan-out (phase 11); other GC requests run on the round thread; `1` runs the covered requests sequentially. `cas_gc_read_concurrency` is rejected without an alias; use `cas_gc_io_concurrency` instead |
785785

786786
The fold-batching controls `gc_fold_threshold` (default 1), `gc_fold_max_defer_rounds` (default 8)
787787
and `gc_frontier_probe_budget` (default unbounded) are internal `PoolConfig` fields with no disk

‎docs/en/antalya/cas/architecture/read-path.md‎

Lines changed: 6 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -12,7 +12,12 @@ doc_type: 'reference'
1212
A `CAS` read never touches a classical local-metadata path: there is no local directory listing to
1313
consult, only a ref resolve followed by object-store reads. This page covers the three ways a file
1414
access is served, the full chain for the common case, the two caches that sit on that chain, and
15-
how a part still open inside a write transaction serves its own reads.
15+
how a part still open inside a write transaction serves its own reads. A directory probe on a path
16+
inside a part (`<table>/<part>/<file>`, which `MergeTree` issues for every checksum entry at load)
17+
is answered from the part's retained folder manifest: a plain file is not a directory, a nested
18+
directory is and lists its children. No object-store `LIST` is involved; only a probe whose part
19+
does not resolve falls back to a listing: the table-level file listing on an `Atomic` table, the
20+
mirrored live-tree listing on a non-`Atomic` table.
1621

1722
## How a file access is served {#access-kinds}
1823

‎docs/en/antalya/cas/configuration.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -106,7 +106,7 @@ entirely before release. Treat this table as a snapshot of the current build, no
106106
| `cas_part_folder_cache_max_entry_bytes` | 16 MiB | Oversized part-folder views bypass retention above this size |
107107
| `cas_manifest_decode_cache_bytes` | 128 MiB | Manifest decode cache byte budget (`0` disables) |
108108
| `cas_gc_meta_pool_size` | `16` | Bounded pool size for GC per-hash freshness-meta writes |
109-
| `cas_gc_read_concurrency` | `16` | Bounded pool size for the GC fold's read-ahead of checkpoints, ref logs, manifests and zero-candidate HEADs; `1` disables |
109+
| `cas_gc_io_concurrency` | `16` | Bounded pool size for GC object-storage requests that run in parallel: the fold's read-ahead (checkpoints, ref logs, manifests, zero-candidate HEADs), the orphan-manifest sweep planning reads, the `SYSTEM CAS GC REBUILD` read-ahead, and the `pending_deletes` blob `HEAD` + conditional `DELETE` fan-out. Not covered: meta writes (`cas_gc_meta_pool_size`) and all other GC requests, which run on the round thread. `1` runs the covered requests sequentially. `cas_gc_read_concurrency` is rejected without an alias; use `cas_gc_io_concurrency` instead |
110110
| `cas_attempt_timeout_ms` | `5000` | Budget for one HTTP attempt of a writable Native mount's control-plane requests (read, head, list, remove, conditional write), at least 1. Together with the connect cap it forms the attempt envelope (`cas_attempt_timeout_ms + 2 × cap`; the cap is `cas_attempt_timeout_ms` itself when the disk's `connect_timeout_ms` is `0`, else `min(connect_timeout_ms, cas_attempt_timeout_ms)`) that the lease arithmetic reserves: one TCP connect and one TLS handshake under the cap each, send/receive bounded per socket operation by `cas_attempt_timeout_ms`. With background renewal the cadence check requires `cas_mount_renew_period_ms + 2 × envelope + cas_lease_safety_margin_ms < cas_mount_lease_ttl_ms`, which puts an effective ceiling on the frozen connect cap: under the defaults (TTL 30000, period 10000, margin 2000) the envelope must stay under 9000, so a disk `connect_timeout_ms` of 2000 ms or more refuses to open writable — lower the connect timeout or raise the TTL if you hit this |
111111
| `cas_lease_safety_margin_ms` | `2000` | Startup-only margin validated against the mount lease TTL: the attempt envelope + `cas_lease_safety_margin_ms` must be strictly less than the mount lease TTL, and `cas_mount_renew_period_ms` + 2 × envelope + `cas_lease_safety_margin_ms` too, or the disk refuses to open writable |
112112
| `cas_unsafe_remount_no_delay` | `0` | Reclaim a mount slot that carries this server's own uuid at once after a hard restart, without observing the slot's token for the lease TTL. Unsafe whenever two processes can hold the same `server_uuid` (a copied uuid file, a stalled predecessor). After such a reclaim the predecessor can still start conditional writes until its own cutoff (`confirmed deadline − cas_lease_safety_margin_ms − 2 × envelope`) or until its next renewal meets the token guard, and a request it already sent may still materialize later. That is not a data hazard: ref-log keys carry `(writer_epoch, sequence)` and creates are conditional, so two writers can never commit different bodies to one key, and recovery's epoch seal settles any straggler (recovery fails closed after 64 successive seal-create attempts displaced by newly materializing old-epoch transactions). The exposure is availability, not data. Intended for test stands and deployments that guarantee one process per uuid |

‎docs/en/antalya/cas/operations/debugging.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -56,7 +56,7 @@ SELECT server_root_id, is_leader, state, last_success_age_seconds, pending_recla
5656
FROM system.cas_mounts WHERE disk = 'cas';
5757

5858
SELECT event_time, outcome, candidates_marked, entries_condemned, entries_graduated,
59-
entries_redeleted, anomalies
59+
entries_redeleted, entries_redelete_failed, anomalies
6060
FROM system.cas_gc_log
6161
WHERE event_type = 'Finish' AND disk_name = 'cas'
6262
ORDER BY event_time DESC LIMIT 10;
@@ -193,8 +193,8 @@ SYSTEM CAS GC RUN cas;
193193

194194
One row per disk it ran on: `disk`, `acquired_lease`, `deferred`, `round`, `candidates_marked`,
195195
`objects_deleted`, `objects_absent`, `objects_replaced`, `objects_spared`, `manifests_deleted`,
196-
`entries_condemned`, `entries_graduated`, `entries_redeleted`, `fence_outs`, `anomalies`,
197-
`pending_candidates`, `pending_condemned`, `pending_retired`. Omitting
196+
`entries_condemned`, `entries_graduated`, `entries_redeleted`, `entries_redelete_failed`, `fence_outs`,
197+
`anomalies`, `pending_candidates`, `pending_condemned`, `pending_retired`. Omitting
198198
the disk name runs one round on every content-addressed disk on the node. A manual run executes
199199
regardless of `SYSTEM CAS GC STOP` — `STOP` pauses only the background scheduler.
200200

‎docs/en/antalya/protocol.md‎

Lines changed: 50 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,50 @@
1+
---
2+
description: 'How the Antalya fork versions its own wire-protocol changes independently of upstream ClickHouse'
3+
sidebar_label: 'Antalya Protocol Version'
4+
sidebar_position: 40
5+
slug: /antalya/protocol
6+
title: 'Antalya Protocol Version'
7+
doc_type: 'reference'
8+
---
9+
10+
# Antalya protocol version {#antalya-protocol-version}
11+
12+
Antalya versions its own wire-protocol changes with `DBMS_ANTALYA_PROTOCOL_VERSION`, a counter that
13+
upstream ClickHouse cannot reach, defined in `src/Core/AntalyaProtocol.h`. A server advertises it in
14+
the `ServerHello` name string, on every connection:
15+
16+
```text
17+
server -> client "ClickHouse (antalya:1)"
18+
```
19+
20+
The client parses the suffix, caps the value with `min(own, server)` and keeps the result. `0` means
21+
the peer is not an Antalya build. Negotiation is per hop and not transitive: initiator to worker and
22+
worker to worker negotiate independently.
23+
24+
Version 1 is the advertisement itself. Nothing is gated on it yet.
25+
26+
## Adding an Antalya-only wire change {#adding-a-wire-change}
27+
28+
- Bump `DBMS_ANTALYA_PROTOCOL_VERSION` by one and gate the change on the negotiated value.
29+
- Never bump `DBMS_TCP_PROTOCOL_VERSION`, and never take a slot in
30+
`DBMS_CLUSTER_PROCESSING_PROTOCOL_VERSION` for a feature upstream does not have.
31+
- Keep the counter cumulative. A backport takes the whole contiguous range up to the value it needs,
32+
or does not bump at all - the `min(own, server)` cap is only sound for a cumulative feature set.
33+
- Gate only what the *client* decides to do. The server never learns the client's version, because
34+
only the server advertises.
35+
- Update this page, and update `docs/en/interfaces/specs/NativeProtocol.md` when the change alters
36+
a packet layout described there.
37+
38+
## Why a counter of our own {#why-a-counter-of-our-own}
39+
40+
An upstream rebase can reuse the next value of an upstream protocol counter for a different feature.
41+
Keeping the Antalya counter separate prevents the same version from describing two wire layouts.
42+
43+
## Why the marker rides in `ServerHello` {#why-the-marker-rides-in-serverhello}
44+
45+
The client `Hello` cannot advertise the version because it is sent before the peer is known. Its
46+
`client_name` is also stored and validated against the Query packet, so changing it can raise
47+
`CLIENT_INFO_DOES_NOT_MATCH` on an upstream peer.
48+
49+
The server advertises through `server_name`, which is display text. The marker stays inside that
50+
existing string because adding a field would make older peers read it as the next packet.

‎docs/en/interfaces/specs/NativeProtocol.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -440,7 +440,7 @@ Server → Client. The reply to ClientHello on successful authentication.
440440

441441
| # | Field | Type | Role | Condition | Description |
442442
|---|------------------|---------|-----------|------------------------|-------------|
443-
| 1 | server_name | String | universal | always | Server identifier |
443+
| 1 | server_name | String | universal | always | Server identifier. An Altinity Antalya build appends `" (antalya:N)"`, where `N` is its Antalya protocol version; a client may ignore or strip the suffix. See [Antalya protocol version](/antalya/protocol). |
444444
| 2 | version_major | VarUInt | universal | always | Server major version |
445445
| 3 | version_minor | VarUInt | universal | always | Server minor version |
446446
| 4 | protocol_version | VarUInt | universal | always | Server's protocol version |

‎docs/en/operations/storing-data.md‎

Lines changed: 14 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -550,9 +550,20 @@ disk-level and server-level settings surface.
550550
- `cas_gc_meta_pool_size` — `16` by default. Bounded thread-pool size for the GC's per-hash freshness-meta
551551
writes (condemn/spare/delete), so a mass `DROP` condemning millions of blobs does not run fully
552552
sequentially.
553-
- `cas_gc_read_concurrency` — `16` by default. Bounded thread-pool size for the GC fold's read-ahead of
554-
checkpoints, ref logs, manifest bodies and zero-candidate `HEAD`s. The fold's decisions stay on the
555-
round thread in their original order; only the fetches overlap. `1` disables read-ahead.
553+
- `cas_gc_io_concurrency` — `16` by default. Bounded thread-pool size for the GC object-storage requests
554+
that run in parallel. It covers:
555+
- the fold's read-ahead of checkpoints, ref logs, manifest bodies and zero-candidate `HEAD`s;
556+
- the planning reads of the orphan-manifest sweep and the read-ahead of `SYSTEM CAS GC REBUILD`;
557+
- the `pending_deletes` phase, which runs one `HEAD` and one conditional `DELETE` (`If-Match`) per blob.
558+
559+
It does not cover the per-hash freshness-meta writes (`cas_gc_meta_pool_size`) or any other GC request
560+
(`LIST`, `gc/state` updates, manifest and ref-object batch deletes, generation pruning, the namespace
561+
janitor, orphan-manifest deletes): those run on the round thread. Decisions, outcomes, events and the
562+
audit log stay on the round thread in their original order; only the requests overlap. An entry is
563+
recorded as deleted only if its own `HEAD` and `DELETE` ran; entries whose request failed, or that were
564+
not submitted, stay pending and are retried in the next round. `1` runs all covered requests
565+
sequentially on the round thread.
566+
`cas_gc_read_concurrency` is rejected without an alias; use `cas_gc_io_concurrency` instead.
556567
- `skip_access_check` — `false` by default. Skips the disk's `CAS` capability probe ("start now,
557568
fix later"). The server-level `skip_access_check` flag skips the generic disk access check;
558569
this disk key governs the `CAS` capability probe.

‎docs/en/operations/system-tables/cas_gc_log.md‎

Lines changed: 6 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -48,6 +48,7 @@ specified (it is enabled by default in the shipped `config.xml`).
4848
- `entries_condemned` ([UInt64](/sql-reference/data-types/int-uint)) — Retired entries newly condemned this round (retired-cursor pipeline stage 1).
4949
- `entries_graduated` ([UInt64](/sql-reference/data-types/int-uint)) — Retired entries newly floor-passed and republished `delete_pending` this round (pipeline stage 2; deleted the next round).
5050
- `entries_redeleted` ([UInt64](/sql-reference/data-types/int-uint)) — Pending exact-token blob deletes executed this round (pipeline stage 3).
51+
- `entries_redelete_failed` ([UInt64](/sql-reference/data-types/int-uint)) — Pending blob deletes whose `HEAD` or exact-token `DELETE` failed this round. Each such entry stays `delete_pending` and is retried in the next round; a non-zero value fails the round.
5152
- `fence_outs` ([UInt64](/sql-reference/data-types/int-uint)) — Expired mounts fenced out by this round's heartbeat floor.
5253
- `anomalies` ([UInt64](/sql-reference/data-types/int-uint)) — Fold clamps surfaced (and survived) this round. A steady non-zero value warrants a look at the round log details.
5354
- `duration_ms` ([UInt64](/sql-reference/data-types/int-uint)) — The round wall-clock duration (on a `Finish` row).
@@ -80,7 +81,7 @@ The phases, in execution order:
8081
| `fold_ref_intake` | Read and fold every new ref log and the manifest bodies its edges name. | one `GET` per new log, one `GET` per manifest edge |
8182
| `fold_reduce` | The per-shard in-degree merge: condemn, spare, graduate. | prior-run streaming `GET`s, one `HEAD` per zero-transition candidate, run `PUT`s |
8283
| `fold_seal_write` | Publish the new fold seal. | one `PUT` |
83-
| `pending_deletes` | The single content-delete site: exact-token deletes of previously published `delete_pending` entries, plus the outcome logs. | one `DELETE` per entry, one outcome-log `PUT` per shard |
84+
| `pending_deletes` | The single content-delete site: exact-token deletes of previously published `delete_pending` entries, plus the outcome logs. `phase_metrics` carries `jobs_scheduled` (entries sent to the GC I/O pool; `0` for no candidates, sequential execution, singleton batches, or refusal before the first submission) and `jobs_failed`. | one `HEAD` and at most one `DELETE` per entry, one outcome-log `PUT` per shard with rows |
8485
| `meta_pool_wait` | Drain the round's per-hash freshness-meta writes. | none on this thread — see the caveat below |
8586
| `round_commit` | The generation-retention prune and the round's single `gc/state` compare-and-swap. | prune `LIST`s and deletes, one compare-and-swap |
8687
| `handoff_reclaim` | Wholesale-reclaim generations a moved run ref stranded below the retention cursor. | prefix `LIST`s and deletes |
@@ -128,6 +129,10 @@ Two caveats when reading these rows:
128129
- Work scheduled onto the GC meta pool runs on other threads, so the `meta_pool_wait` row's
129130
`ProfileEvents` delta is **empty by construction**. Read its `phase_metrics` `jobs_scheduled` /
130131
`jobs_completed` next to its duration instead: they distinguish a deep queue from a slow endpoint.
132+
- Requests issued by GC I/O pool workers — the `pending_deletes` fan-out and the fold read-ahead of
133+
`fold_ref_intake` and `fold_reduce` — also run on other threads and are missing from those phase
134+
rows' `ProfileEvents`. They still count in `system.events`. On `pending_deletes`, read
135+
`phase_metrics` `jobs_scheduled` / `jobs_failed` next to `redeleted`.
131136
- Phase durations do not sum to the round's `duration_ms`. The round also performs untimed
132137
bookkeeping between phases, and the `Finish` row's `duration_ms` remains the authority on total
133138
round time.

‎docs/en/sql-reference/statements/system.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -479,7 +479,7 @@ When `disk_name` is given, the round runs on that content-addressed disk only; t
479479

480480
Each round is recorded in [`system.cas_gc_log`](/operations/system-tables/cas_gc_log) as a `Start` and a `Finish` row (with `trigger = 'Manual'`).
481481

482-
The command returns one row per disk it ran on (multiple rows when `disk_name` is omitted), with columns `disk`, `acquired_lease`, `deferred`, `round`, `candidates_marked`, `objects_deleted`, `objects_absent`, `objects_replaced`, `objects_spared`, `manifests_deleted`, `entries_condemned`, `entries_graduated`, `entries_redeleted`, `fence_outs`, `anomalies`, `pending_candidates`, `pending_condemned`, and `pending_retired`, describing the outcome of that round. The `pending_*` columns are the retire pipeline's remaining backlog sizes read from the `gc/state` this round's own commit just published (not this round's own delta, unlike the columns before them) — `0` on a non-authoritative row (`acquired_lease = 0` or `deferred = 1`), same as every other counter.
482+
The command returns one row per disk it ran on (multiple rows when `disk_name` is omitted), with columns `disk`, `acquired_lease`, `deferred`, `round`, `candidates_marked`, `objects_deleted`, `objects_absent`, `objects_replaced`, `objects_spared`, `manifests_deleted`, `entries_condemned`, `entries_graduated`, `entries_redeleted`, `entries_redelete_failed`, `fence_outs`, `anomalies`, `pending_candidates`, `pending_condemned`, and `pending_retired`, describing the outcome of that round. The `pending_*` columns are the retire pipeline's remaining backlog sizes read from the `gc/state` this round's own commit just published (not this round's own delta, unlike the columns before them) — `0` on a non-authoritative row (`acquired_lease = 0` or `deferred = 1`), same as every other counter.
483483
484484
A manual run always executes, regardless of [`SYSTEM CAS GC STOP`](#system-cas-gc-stop-start): `STOP` pauses only the background scheduler on that disk.
485485

0 commit comments

Comments
 (0)