Skip to content

swarms: deterministic join/union scenario names and seeded sampling - #171

Open
CarlosFelipeOR wants to merge 3 commits into
mainfrom
swarms-deterministic-test-names
Open

CarlosFelipeOR wants to merge 3 commits into
mainfrom
swarms-deterministic-test-names

Conversation

@CarlosFelipeOR

@CarlosFelipeOR CarlosFelipeOR commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

/swarms/feature/swarm joins/join clause/* and /swarms/feature/swarm union/union clause/* built the scenario name from the tables' SQL. Iceberg and MergeTree tables are named with getuid(), so the names changed on every run:

join 996 of 816480: s3('http://minio:9000/...', 'admin', 'password') with database_d02f8ef1_b498_11f1_....`namespace_..._table_...` in allow mode ...

The number was the position in the sample, and the sample was fixed by random.seed(42), so every run executed the same 1000 of the 816480 join combinations (mostly the same 500 union combinations). Even so, the same join N has pointed to three different combinations since July, because the global RNG is also consumed elsewhere (every index moved on 2026-08-18).

Change

  • A scenario is named after its combination: join N of <total>: <parameters>, where N is the combination's position in the full product. The same name is always the same combination, in any run.
  • Each run executes a different sample (1000 join, 500 union), so the numbers of a run are not sequential. Scenarios run in ascending order.
  • The seed is --seed, else GITHUB_RUN_ID, else random, and is printed at the start of the feature ([note] join combinations seed: <seed>). All jobs of a CI run, and a rerun of any of them, execute the same combinations; different runs cover different combinations.
  • The sample uses its own random.Random(seed), so other random calls can no longer shift it.
  • Table labels (iceberg_table_data1, s3Cluster_data2, ...) replace the SQL in the name. The join condition is shown without the t1./t2. prefixes.
  • Union: UNION ALL/UNION DISTINCT and the object_storage_cluster setting of each side (None = unset) are now part of the combination instead of being drawn inside the parallel scenarios. Total is 31752 (was 11664 plus random choices). Each side is now unset in 1 of 7 combinations instead of 1 of 2, and combinations whose cluster was not used are gone.
  • Removed the join 455 of 816480* xfail (#1244): since 2026-08-18 that index was a different combination, so it already did not cover the issue. It is not replaced in this PR.

Examples:

join 963 of 816480: iceberg_table_data1 with iceberg_table_data1 in allow mode on replicated_cluster_three_nodes cluster with ASOF JOIN on timestamp_col = timestamp_col
union 41 of 31752: iceberg_table_data1 UNION ALL iceberg_table_data1 in cluster_one_observer_one_swarm cluster and cluster_one_observer_two_swarm cluster

Reproducing a failure

Use the seed of the failed run (its GITHUB_RUN_ID, also printed in the log); otherwise join N is probably not in the sample:

./regression.py ... --seed <seed> --only "/swarms/feature/swarm joins/join clause/join N of*"

Things to be aware of

  • A combination only appears in the runs that sample it, so its history in clickhouse_regression_results is sparse.
  • The seed is only in the log, not in the database (known limitation); in CI it is the run id in job_url.
  • Table data (swarm_steps) still uses the global random.seed(42), unchanged.

Verification (local, antalya 26.6.4, --with-analyzer)

Suite Seed source Seed Result Same scenarios as
join random 2638357907 1000/1000 OK -
join GITHUB_RUN_ID 2638357907 1000/1000 OK run above
join --seed 2638357907 1000/1000 OK runs above
join GITHUB_RUN_ID (real run id) 36275169144 1000/1000 OK different sample, 1 id in common
union GITHUB_RUN_ID 36275169144 500/500 OK -
union --seed 36275169144 500/500 OK run above

In every run, each join N / union N received exactly combination N of the full product, its name matched its arguments, and the ids matched the sample recomputed offline from the seed. For union, the executed queries carried exactly the expected object_storage_cluster settings.

🤖 Generated with Claude Code

CarlosFelipeOR and others added 3 commits September 23, 2026 19:53
Scenario names embedded the tables' SQL, which carries getuid() table
names, so every run produced different names. Names are now just
"join N of 1000" / "union N of 500"; the parameters are in the
scenario's arguments in the log.

The combination sample is drawn from a per-run seed (random, or
--seed) with its own random.Random, so each run covers different
combinations and any run can be reproduced. The union per-scenario
choices (UNION ALL/DISTINCT, object_storage_cluster settings) are
drawn in the feature instead of inside the parallel scenarios.

Remove the "join 455 of 816480" xfail for #1244: it had already
stopped matching the intended combination.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…N_ID

"join N of <total>" is now the combination's position in the full
product plus its parameters, so the same name is always the same
combination. The seed defaults to GITHUB_RUN_ID, so all jobs of a CI
run and their reruns sample the same combinations. Union ALL/DISTINCT
and object_storage_cluster (None = unset) are now part of the combination.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant