Skip to content

Antalya 26.6: Allow empty object storage cluster - #2221

Open
ianton-ru wants to merge 23 commits into
antalya-26.6from
feature/antalya-26.6/object_storage_cluster_allow_empty
Open

ianton-ru wants to merge 23 commits into
antalya-26.6from
feature/antalya-26.6/object_storage_cluster_allow_empty

Conversation

@ianton-ru

Copy link
Copy Markdown

Rebase of #2028

Changelog category (leave one):

  • Improvement

Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md):

Allow empty object storage cluster

Documentation entry for user-facing changes

With 'object_storage_cluster' setting query to s3,iceberg and some other sources are executed as cluster request.
But with swarm cluster, when initiator is not a cluster member, may be situation when no one swarm node is alive at the moment. In this case query is failed with CLUSTER_DOESNT_EXIST error.

New setting object_storage_cluster_fallback_if_empty allow to execute read query on local node in this case.

Write query is not executed on cluster right now, so attempt to write is still failed in this case to avoid situation when query is success when swarm is empty and failed when has some nodes alive.

PR is a little bit complex because:
s3(...) - can fall back if object_storage_cluster is empty (cluster does not have active nodes, not 'empty setting value')
s3(...) SETTINGS object_storage_remote_initiator=1 - failed on local node if object_storage_cluster is empty
s3(...) SETTINGS object_storage_remote_initiator=1, object_storage_remote_initiator_cluster='...' - decision about falling back must be made on remote initiator, on local node object_storage_cluster can be unknown.

But behavior is not changed for Cluster functions:

s3Cluster(...) - can't fall back
s3Cluster(...) SETTINGS object_storage_remote_initiator=1 - must failed on remote initiator if object_storage_cluster is empty.

CI/CD Options

Exclude tests:

  • Fast test
  • Integration Tests
  • Stateless tests
  • Stateful tests
  • Performance tests
  • Aarch64 tests
  • All with ASAN
  • All with TSAN
  • All with MSAN
  • All with UBSAN
  • All with Coverage
  • All Regression
  • Disable CI Cache

Regression jobs to run:

  • Fast suites (mostly <1h)
  • Aggregate Functions (2h)
  • Alter (1.5h)
  • Benchmark (30m)
  • ClickHouse Keeper (1h)
  • Iceberg (2h)
  • LDAP (1h)
  • OAuth (5m)
  • Parquet (1.5h)
  • RBAC (1.5h)
  • SSL Server (1h)
  • S3 (2h)
  • S3 Export (2h)
  • Swarms (30m)
  • Tiered Storage (2h)

ianton-ru and others added 20 commits August 17, 2026 13:50
Share cluster resolution via resolveClusterRead so getQueryProcessingStage matches read when object_storage_cluster_fallback_if_empty is enabled, and skip local object_storage_cluster lookup when object_storage_remote_initiator and object_storage_remote_initiator_cluster are both set.

Co-authored-by: Cursor <cursoragent@cursor.com>
Cover pure fallback on unknown cluster, aggregate planning, remote-initiator interaction, and integration scenarios with locally unknown object_storage_cluster.

Co-authored-by: Cursor <cursoragent@cursor.com>
Cover stateless and integration cases where object_storage_cluster is missing locally and remote initiator falls back to non-cluster execution.

Co-authored-by: Cursor <cursoragent@cursor.com>
…uster functions.

Distinguish s3() with object_storage_cluster setting from s3Cluster() argument
so fallback works for the former but explicit cluster names still fail when unknown.

Co-authored-by: Cursor <cursoragent@cursor.com>
Distinguish alternative-syntax table functions from table reads so remote
initiator keeps s3() for fallback and uses cluster path for Iceberg tables.

Co-authored-by: Cursor <cursoragent@cursor.com>
… cluster setting.

Track explicit *Cluster function arguments separately from table engine and
query object_storage_cluster settings so fallback works for persistent tables.

Co-authored-by: Cursor <cursoragent@cursor.com>
Split storage policy from the setting check and route all local-fallback decisions through one helper so the resolve/read path is easier to follow.

Co-authored-by: Cursor <cursoragent@cursor.com>
… syntax.

Pure-send to a remote initiator requires object_storage_remote_initiator_cluster; otherwise keep the clustered path that defaults the initiator cluster to object_storage_cluster.

Co-authored-by: Cursor <cursoragent@cursor.com>
…ster.

Under remote-initiator deferral, an empty local cluster name must fall back to pure send so ENGINE=S3/Iceberg without object_storage_cluster does not hit LOGICAL_ERROR.

Co-authored-by: Cursor <cursoragent@cursor.com>
Drop the redundant pure-send gate, flatten fallback_to_pure, and reuse the already resolved remote-initiator cluster on the clustered path.

Co-authored-by: Cursor <cursoragent@cursor.com>
Local fallback applies only to object_storage_cluster; a bad remote-initiator cluster must still report CLUSTER_DOESNT_EXIST.

Co-authored-by: Cursor <cursoragent@cursor.com>
Empty OSC with remote_initiator and no remote_initiator_cluster must keep BAD_ARGUMENTS even when fallback is enabled; extend tests for cases 1-3 and this regression.

Co-authored-by: Cursor <cursoragent@cursor.com>
Pure-send ENGINE tables under remote-initiator deferral (preserving OSC in SETTINGS/context) so fallback matches s3() alternative syntax; extend TF and ENGINE tests for cases 1-3.

Co-authored-by: Cursor <cursoragent@cursor.com>
Merge the two policy hooks into usesObjectStorageClusterSettingSyntax, share local-vs-remote fallback decisions between read and getQueryProcessingStage, and dedupe remote-initiator send.

Co-authored-by: Cursor <cursoragent@cursor.com>
…TTINGS.

object_storage_cluster may be unknown locally and defined on the remote (or the reverse); *Cluster would bake the name into the function argument and skip remote fallback.

Co-authored-by: Cursor <cursoragent@cursor.com>
…cluster.

Align flag semantics with s3()/iceberg() alternative syntax when the cluster name comes from SETTINGS rather than a *Cluster argument.

Co-authored-by: Cursor <cursoragent@cursor.com>
…ter.

AST rewrite and setting-syntax checks only need cluster_name_from_function_argument.

Co-authored-by: Cursor <cursoragent@cursor.com>
@github-actions

github-actions Bot commented Aug 17, 2026 •

Copy link
Copy Markdown

Workflow [PR], commit [ada4a56]

@ianton-ru
ianton-ru marked this pull request as ready for review August 17, 2026 16:59
@ianton-ru ianton-ru closed this Aug 18, 2026
@ianton-ru ianton-ru reopened this Aug 18, 2026
@ianton-ru ianton-ru mentioned this pull request Aug 18, 2026
11 of 28 tasks
@ianton-ru ianton-ru added antalya port-antalya PRs to be ported to all new Antalya releases antalya-26.6 labels Aug 18, 2026
@ianton-ru

Copy link
Copy Markdown
Author

Failed stateless tests is flaky.

@arthurpassos arthurpassos left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

to be fairly honest, I couldn't do an useful review on this one in a timely manner. Too many branching in code I am not super familiar with. I'll just delay you more if I try to understand every bit of it.

AI review below:

Summary
The change is a gated fail-open for reads when object_storage_cluster names a missing or 0-node cluster. Writes and *Cluster stay fail-closed. That contract is internally consistent.

resolveClusterRead + the RI “send plain s3() + setting” path is the right design for “OSC exists only on the remote.” Tests lock the main s3() / S3 ENGINE / RI cases.

Findings
⚠️ Parallel replicas can be treated as OSC fallback
TableFunctionObjectStorage.cpp builds StorageObjectStorageCluster from cluster_for_parallel_replicas when OSC is unset. The PR defaults cluster_name_from_function_argument = false, so usesObjectStorageClusterSettingSyntax() is true.

Then getClusterName falls through to that constructor name. If fallback is on and the parallel-replicas cluster is unknown/empty, getClusterImpl(..., allow_null=true) turns a failed PR plan into a silent local s3() read.

That is realistic: the people who will put fallback in a profile are the same people who already set cluster_for_parallel_replicas.

Fix: do not apply OSC fallback unless the name actually came from object_storage_cluster (query/table/database). Parallel-replicas construction should look like a function-argument cluster (setClusterNameFromFunctionArgument(true) or a third origin).

⚠️ “Empty” is not “swarm is down”
The PR story is “no swarm node is alive.” The code checks tryGetCluster and getAllNodeCount() == 0. getAllNodeCount is configured replicas (per_replica_pools.size()), not live ones.

A named swarm with every host dead still has count > 0 → no fallback → ALL_CONNECTION_TRIES_FAILED / skip_unavailable_shards. Tests only use unknown names.

Either document that, or the stated swarm case is not solved.

⚠️ getCluster() ignores the new resolver
read() / getQueryProcessingStage use resolveClusterRead. Public getCluster() still uses the constructor name and always throws. INSERT ... SELECT with parallel_distributed_insert_select can throw while the matching SELECT falls back. Fail-closed is defensible; the dual API is not.

💡 PR text uses the wrong setting name
Body: object_storage_cluster_fallback_if_empty. Code/tests: object_storage_cluster_fallback_to_local_if_empty. No alias.

Tests
Good for unknown-name s3() / ENGINE / RI / “don’t mask a missing RI-cluster.” Missing: write still fails, 0-node configured cluster, iceberg/DataLake, and the parallel-replicas collision above.

Verdict
Request changes if this is meant to be a profile default for swarms.

Minimum:

Exclude cluster_for_parallel_replicas from fallback.
Fix the changelog/setting name in the PR body.
Say “unknown or zero configured nodes,” not “swarm is down,” unless you add a liveness check (I would not).

@DimensionWieldr

DimensionWieldr commented Sep 8, 2026 •

Copy link
Copy Markdown
Collaborator

CI triage

No PR-caused failures. The only red product test is pre-existing:

  • 02265_column_ttl on Stateless tests (amd_binary, cas s3 storage, parallel) — CAS fetch/relink race (NETWORK_ERROR on part relink). Same job already fails this test on antalya-26.6 MasterCI (2/2 runs that ran it). In-job diagnosis: 25 fail / 1 pass. Unrelated to this diff. The PR’s 04303_object_storage_cluster_fallback_to_local_if_empty passed in that job.

Grype alpine OpenSSL CVEs and the cancelled OAuth / GrypeScanServer jobs are infrastructure.

Swarms coverage

Not sufficient for the new fallback. The suite never sets object_storage_cluster_fallback_to_local_if_empty. Closest coverage is still the old path: observer-only swarm → CLUSTER_DOESNT_EXIST (cluster_with_only_one_observer_node). Node-failure tests kill or overload workers while others stay up; they do not cover “swarm exists but every worker is dead → read locally”. Writes-must-still-fail and s3Cluster must-not-fall-back are also absent from swarms.

That matrix lives in 04303 (passed here) and test_s3_cluster (not in this PR’s job set). Swarm regression going green (1521/1521) only shows we did not break existing swarm behavior.

I'm going to add some swarms tests.

@DimensionWieldr

Copy link
Copy Markdown
Collaborator

Tests against this PR added here to swarms suite: https://github.com/Altinity/clickhouse-regression/blob/main/swarms/tests/fallback_to_local_if_empty.py

Tests are passing.

LGTM

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

antalya antalya-26.6 port-antalya PRs to be ported to all new Antalya releases verified Approved for release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants