Skip to content

CAS: remove a part with one ref drop - #2452

Merged
filimonov merged 1 commit into
antalya-26.6from
fix/antalya-26.6/cas-part-removal-one-drop-pr
Oct 1, 2026
Merged

filimonov merged 1 commit into
antalya-26.6from
fix/antalya-26.6/cas-part-removal-one-drop-pr

Conversation

@filimonov

Copy link
Copy Markdown
Member

Removing a part from a cas disk now writes one ref-log record and no manifest. Before, every removed part cost two extra manifest writes and five extra ref-log records, more with projections.

DataPartStorageOnDiskBase::remove renames a part to delete_tmp_<part>, unlinks its files and removes the directory in separate disk transactions. On a cas disk a part is one ref, so the rename and the unlink batch each published a new manifest before the final drop. Dropping a ref is atomic, so the rename has nothing to protect: for content-addressed disks remove calls removeSharedRecursive on the part directory and returns. Detached parts keep the generic path.

Measured

Create/insert/drop cycle: CREATE TABLE, 3 inserts of 1000 rows, SELECT, DROP TABLE SYNC; 8 workers, 48 cycles per run, RustFS, two runs per side. Base is antalya-26.6 at e2dd0f5.

Per cycle Base This PR Change
PUT, cas disk 64 28 −56%
GET, cas disk 83 47 −43%
Wall time, cas disk, first pass 30.3 ms 17.0 ms −44%
Wall time, cas disk, second pass 27.9 ms 15.7 ms −44%
PUT, cas / s3 disk 2.08 0.92
GET, cas / s3 disk 13.85 7.85
Wall time, cas / s3 disk, first pass 1.70 0.84

Request counts of workloads that remove no parts in the measured window (inserts, attach, selects, mutations) are the same on both sides.

Events in system.cas_log per removed part:

Event Base This PR
manifest_put 2 (3 with a projection) 0
ref_repoint 1 (2 with a projection) 0
ref_drop 2 1

Behaviour

  • A part being removed from a cas disk keeps its name until its ref is dropped; no delete_tmp_<part> ref appears.
  • The removal of one part is one atomic ref-log record: after a failure the part is either fully present or fully gone.

Related: #2429

Changelog category (leave one):

  • Performance Improvement

Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md):

Removing a data part from a cas disk writes one ref-log record instead of two manifests and six ref-log records. A create/insert/drop cycle issues 56% fewer PUT and 43% fewer GET requests and takes 44% less time.

Documentation entry for user-facing changes

  • Documentation is written (mandatory for new features)

CI/CD Options

Exclude tests:

  • Fast test
  • Integration Tests
  • Stateless tests
  • Stateful tests
  • Unit tests
  • Performance tests
  • Aarch64 tests
  • All with ASAN
  • All with TSAN
  • All with MSAN
  • All with UBSAN
  • All with Coverage
  • All Regression
  • Disable CI Cache

Regression jobs to run:

  • Fast suites (mostly <1h)
  • Aggregate Functions (2h)
  • Alter (1.5h)
  • Benchmark (30m)
  • CAS (content-addressed storage; Antalya only)
  • ClickHouse Keeper (1h)
  • Iceberg (2h)
  • LDAP (1h)
  • OAuth (5m)
  • Parquet (1.5h)
  • RBAC (1.5h)
  • SSL Server (1h)
  • S3 (2h)
  • S3 Export (2h)
  • Swarms (30m)
  • Tiered Storage (2h)

🤖 Generated with Claude Code

`DataPartStorageOnDiskBase::remove` renames a part to `delete_tmp_<part>`,
unlinks its files and removes the directory in separate disk transactions.
On a `cas` disk a part is one ref, and the rename and the unlink batch each
published a new manifest before the final drop: two extra manifest writes
and five extra ref-log records per removed part, more with projections.

Dropping a ref is atomic, so a `cas` disk does not need the rename.
`remove` now calls `removeSharedRecursive` on the part directory for
content-addressed disks: one ref-log record, no manifest write. Detached
parts keep the generic path.

Create/insert/drop cycle (3 parts per table, 8 workers, RustFS), per cycle
on a `cas` disk: PUT 64 -> 28, GET 83 -> 47, wall time 30.3 -> 17.0 ms.

Signed-off-by: Mikhail Filimonov <mfilimonov@altinity.com>
@github-actions

Copy link
Copy Markdown

Workflow [PR], commit [fd7be67]

@filimonov filimonov added the CAS label Sep 30, 2026
@filimonov filimonov mentioned this pull request Oct 1, 2026
68 tasks
@filimonov
filimonov requested a review from k-morozov October 1, 2026 14:09
@filimonov
filimonov merged commit 41824f9 into antalya-26.6 Oct 1, 2026
360 of 371 checks passed
@DimensionWieldr

Copy link
Copy Markdown
Collaborator

PR #2452 CI Triage

Five jobs failed. Six other red statuses are the report links for those same jobs. The pull request is already merged.

Summary

Category Count What
pre-existing-flaky 1 job (8 scenarios, one cause) cas_selects — mount lease dropped during concurrent SELECT FINAL
cascade 7 paths Parent rows of the two regression jobs, behind the leaves below
infrastructure 2 jobs Grype on the Alpine server image and on Keeper: CVE-2026-85091 in zlib 1.3.2-r0
unknown 2 One lightweight-delete scenario, and the ASAN CAS stress hung check

Nothing in the diff is a demonstrated regression. DataPartStorageOnDiskBase::remove is the only product change, and the failures that have a known cause do not go through it.

pre-existing-flaky

/selects/final/force/concurrent/* on RegressionTestsRelease / CAS (selects) / cas_selects. Eight leaves failed with Code: 210, content-addressed disk 'cas_disk' -- mount lease not held. The next sibling, SELECT WHERE parallel inserts, deletes, updates, passed, so the pool came back to Live in the same feature. Four parent paths (/selects through /selects/final/force/concurrent) are cascade from that outage.

This is issue #2332, opened 2026-09-09 on PR #2300, three weeks before this run. A note there records cas_selects still failing 15 of 47 runs after the keep-alive workaround. This pull request is 1 of 1, which that rate predicts. The failing statements are reads. The change in this pull request is the part-removal path.

infrastructure

GrypeScanServer (-alpine) and GrypeScanKeeper both fail /docker vulnerabilities/CVE-2026-85091@nvd:cpe,High. The package is Alpine 3.24.2 zlib 1.3.2-r0 (heap overflow in gz_vacate). The non-Alpine server scan on this same SHA passed. PR #2451, merged the same day with an unrelated diff, fails the same two jobs on the same zlib CVE. Sample is 2 pull requests. The scan will keep failing until the Alpine package moves.

unknown

/lightweight delete/concurrent delete/MergeTree/concurrent delete without overlap with alter delete on cas_lightweight_delete_4. After a lightweight DELETE of the even rows and a concurrent ALTER DELETE of the odd rows, both with mutations_sync=2, SELECT count(*) was 50000 instead of 0. That is one partition's half (10 partitions × 100000 rows). The same scenario passed on the six other engines in this job, and cas_s3_cache_lightweight_delete_4 passed the same filter on this SHA. Issue #2331 is the same race family on a different scenario (random delete entire table without overlap, leftover 1), documented at 5 of 322 runs, and that scenario passed here. One run cannot say whether this leaf fails at that rate or only on this diff.

What would resolve it: the 60-day fail rate of this exact scenario on cas_lightweight_delete_4, branch runs (pull_request_number = 0) against this pull request.

Hung check failed, possible deadlock found on Stress test (amd_asan_ubsan, cas s3 storage). INSERT INTO tab_00625 from 00625_summing_merge_tree_merge.sql was still in the process list after 1644s, already cancelled. The stack is ThreadFuzzer inside S3ObjectStorage::isReadOnly, called from DataPartStorageOnDiskBase::isCaseInsensitive while writing a wide part. No <Fatal> line. stderr.log has a LeakSanitizer report in QueryTreeBuilder / Context, and that report is not the recorded failure and does not touch remove. On this SHA, Stress test (amd_debug, cas s3 storage) and Stress test (amd_asan_ubsan) passed.

What would resolve it: the 60-day fail rate of Hung check failed, possible deadlock found on Stress test (amd_asan_ubsan, cas s3 storage), branch against this pull request.

Recommendations

  1. Treat cas_selects as CAS: concurrent SELECT FINAL drops the mount lease with no object-store outage (cas_selects) #2332. It does not need a change in this pull request.
  2. Bump Alpine zlib past 1.3.2-r0. The Grype failure is the base image, and it is already red on other pull requests.
  3. Close the two unknown rows with the rate queries above before treating either as a defect in the one-ref-drop removal.

@DimensionWieldr

Copy link
Copy Markdown
Collaborator

Looked into the unknown section fails:

Lightweight delete fail is ClickHouse#122083.

ASAN CAS stress fails are not related to this PR. The hung check is a single cancelled INSERT in the ASAN CAS stress job, and that job passed on the MasterCI run of the merge commit.

LGTM

@DimensionWieldr DimensionWieldr added the verified Approved for release label Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants