Affected tests:
/swarms/feature/swarm joins/join clause/join N of 816480: ... — 109 scenarios, all in allow mode (49 expecting exitcode 60, 60 expecting 81)
Affected files:
Description
The swarms suite went from green to 109 failures on antalya-25.8 when Altinity/ClickHouse#2287 bumped 25.8.30 → 25.8.33. Two assertion signatures, both in check_join:
| # |
assertion |
expected |
got |
| 49 |
assert r.exitcode == exitcode |
60 UNKNOWN_TABLE |
0 + rows |
| 60 |
assert result.exitcode in (81, 10) |
81 UNKNOWN_DATABASE |
0 + rows |
Not a product regression. The tests required an error that no longer happens.
Analysis
A JOIN whose leftmost table is an object storage cluster table used to be shipped whole to the swarm nodes, which do not have the other table. ClickHouse#94748 fixed it by wrapping the left table in a column-reading subquery — the same treatment non-leftmost remote tables already got via is_remote. It reached 25.8 through backport ClickHouse#114621, which the bump pulled in.
joins.py had that failure hard-coded as expected under check_clickhouse_version("<26.3"), in 31 places.
The new behaviour is correct:
- For all 109 failing scenarios, the query output is identical to the suite's own reference in
get expected result from joining merge tree tables (the same JOIN over plain MergeTree copies) — 109/109.
- Across all 1000 leaf join scenarios in the run, zero produced a
DB::Exception. The expected-error branch was dead code on 25.8.33.
Why <26.3 worked until now, and why it broke:
| branch |
version |
fix present |
gate expects |
result |
| antalya-25.8 (before) |
25.8.30 |
no |
error |
pass |
| antalya-25.8 (bump) |
25.8.33 |
yes |
error |
109 fail |
| antalya-26.1 |
26.1.11 |
no |
error |
pass |
| antalya-26.6 |
26.6.2 |
yes |
success |
pass |
The gate assumed the fix only existed from 26.3 up. #94748 first shipped upstream in 26.2, not 26.1: it was merged 2026-01-26 20:29 UTC and the 26.1 release branch was cut at 18:20 the same day — it missed by 2h09m. No 26.1 tag contains it (verified on v26.1.1.912, v26.1.8.35, v26.1.12.23). The Merged into: 26.1.1.899 line in the PR body is the master build number at merge time and is misleading here.
Solution
Replaced the 31 version checks with a named predicate matching where the fix actually landed:
cluster_join_fix_missing = check_clickhouse_version("~26.1")(self) or check_clickhouse_version("<25.8.33")(self)
| version |
<26.3 (old) |
predicate |
correct |
| 25.8.30 |
True |
True |
✓ |
| 25.8.33 |
True ✗ |
False |
✓ |
| 26.1.11 |
True |
True |
✓ |
| 26.6.2 |
False |
False |
✓ |
Three commits, because the first swap was too broad:
0dd36abd7 — introduce the predicate, replace all 31 gates.
d019c5637 — restore the FULL OUTER JOIN xfail gate at the top of check_join. It is evaluated before the predicate is defined, so the first commit raised UnboundLocalError on 31 scenarios. It is also a different facet of ClickHouse#89996 that #94748 did not fix, and since xfail() aborts immediately (measured: <17 ms) those scenarios never ran the query — there is no evidence the case works on 25.8.33.
302ae29c7 — restore the 20 gates of the # resulting rows are not stable block. That block only tolerates unstable rows and can never cause a failure; turning it off demanded exact data equality in 20 combinations that #94748 does not address.
Only the two expected-error blocks needed the new predicate: exitcode = 81 (5 gates, lines 204-246) and exitcode = 60 (5 gates, lines 249-295).
Result on antalya-25.8 after all three: 1488 ok, 32 xfail, 0 join failures (was 1380 ok / 109 fail / 32 xfail). The 32 xfail are the pre-existing FULL OUTER JOIN is not supported in allow mode plus one unrelated task-rescheduling xfail.
Follow-ups (not addressed)
- The predicate hard-codes 26.1 as missing the fix. If #94748 is ever backported there, the gate breaks in the same direction this issue describes.
- Running the suite without
--with-analyzer produces spurious failures. join_conditions includes inequality conditions and product() pairs them with every join clause; non-ASOF joins reject them with Code: 403 INVALID_JOIN_ON_EXPRESSION, and it fires on the reference MergeTree query in check_join. CI passes --with-analyzer (regression.yml:474), so it is only visible locally — 45 failures in one filtered local run.
References
Affected tests:
/swarms/feature/swarm joins/join clause/join N of 816480: ...— 109 scenarios, all inallow mode(49 expecting exitcode 60, 60 expecting 81)Affected files:
swarms/tests/joins.pyDescription
The
swarmssuite went from green to 109 failures onantalya-25.8when Altinity/ClickHouse#2287 bumped 25.8.30 → 25.8.33. Two assertion signatures, both incheck_join:assert r.exitcode == exitcode60UNKNOWN_TABLE0+ rowsassert result.exitcode in (81, 10)81UNKNOWN_DATABASE0+ rowsNot a product regression. The tests required an error that no longer happens.
Analysis
A JOIN whose leftmost table is an object storage cluster table used to be shipped whole to the swarm nodes, which do not have the other table. ClickHouse#94748 fixed it by wrapping the left table in a column-reading subquery — the same treatment non-leftmost remote tables already got via
is_remote. It reached 25.8 through backport ClickHouse#114621, which the bump pulled in.joins.pyhad that failure hard-coded as expected undercheck_clickhouse_version("<26.3"), in 31 places.The new behaviour is correct:
get expected result from joining merge tree tables(the same JOIN over plain MergeTree copies) — 109/109.DB::Exception. The expected-error branch was dead code on 25.8.33.Why
<26.3worked until now, and why it broke:The gate assumed the fix only existed from 26.3 up. #94748 first shipped upstream in 26.2, not 26.1: it was merged 2026-01-26 20:29 UTC and the 26.1 release branch was cut at 18:20 the same day — it missed by 2h09m. No 26.1 tag contains it (verified on
v26.1.1.912,v26.1.8.35,v26.1.12.23). TheMerged into: 26.1.1.899line in the PR body is the master build number at merge time and is misleading here.Solution
Replaced the 31 version checks with a named predicate matching where the fix actually landed:
<26.3(old)Three commits, because the first swap was too broad:
0dd36abd7— introduce the predicate, replace all 31 gates.d019c5637— restore theFULL OUTER JOINxfailgate at the top ofcheck_join. It is evaluated before the predicate is defined, so the first commit raisedUnboundLocalErroron 31 scenarios. It is also a different facet of ClickHouse#89996 that #94748 did not fix, and sincexfail()aborts immediately (measured: <17 ms) those scenarios never ran the query — there is no evidence the case works on 25.8.33.302ae29c7— restore the 20 gates of the# resulting rows are not stableblock. That block only tolerates unstable rows and can never cause a failure; turning it off demanded exact data equality in 20 combinations that #94748 does not address.Only the two expected-error blocks needed the new predicate:
exitcode = 81(5 gates, lines 204-246) andexitcode = 60(5 gates, lines 249-295).Result on
antalya-25.8after all three: 1488 ok, 32 xfail, 0 join failures (was 1380 ok / 109 fail / 32 xfail). The 32 xfail are the pre-existingFULL OUTER JOIN is not supported in allow modeplus one unrelated task-rescheduling xfail.Follow-ups (not addressed)
--with-analyzerproduces spurious failures.join_conditionsincludes inequality conditions andproduct()pairs them with every join clause; non-ASOF joins reject them withCode: 403 INVALID_JOIN_ON_EXPRESSION, and it fires on the reference MergeTree query incheck_join. CI passes--with-analyzer(regression.yml:474), so it is only visible locally — 45 failures in one filtered local run.References
FULL OUTER JOINworks incorrectly with Cluster functions ClickHouse/ClickHouse#89996-Clustertable functions ClickHouse/ClickHouse#94748 (26.2)-Clustertable functions ClickHouse/ClickHouse#114621 (25.8.33)