Skip to content

BACKUP of CAS tables to S3 on the same endpoint - #2415

Open
k-morozov wants to merge 21 commits into
antalya-26.6from
cas/backup-native-copy-bug
Open

k-morozov wants to merge 21 commits into
antalya-26.6from
cas/backup-native-copy-bug

Conversation

@k-morozov

@k-morozov k-morozov commented Sep 22, 2026 •

Copy link
Copy Markdown

BACKUP of a table on a CAS to an S3 destination that shares the pool's endpoint failed outright, and the code path that failed would have silently corrupted the backup had it not failed.

BackupWriterS3::copyFileFromDisk asks S3 to copy the source object server-side. That assumes the file is that object, whole, from byte 0. On a CAS disk neither half holds:

  • A small per-part file (checksums.txt, count.txt, columns.txt, ...) is an inline manifest entry with no object of its own. getStorageObjects returns a sized placeholder with an empty remote key, deliberately poisoned so a reader that bypasses the read path fails loudly rather than reading someone else's bytes. The copy passed that empty key to CopyObject, so every BACKUP died with Invalid argument - these files exist in every part.

  • A blob object is [envelope][payload] and the payload is the file. BlobLocation carries the payload offset, but getStorageObjects returns only key and length and getBlobPath keeps only the key, so the copy started at byte 0 and pulled the envelope into the backup. RESTORE then failed.

    The first failure masked the second: checksums.txt aborted the backup before any blob corruption
    could surface. Fixing only the empty key would have turned a loud failure into backups that report
    success and cannot be restored.

The same bug was on the BACKUP ... TO Disk(...) path (an s3/s3_plain disk on the pool's endpoint): DiskObjectStorage::copyFile -> copyFileImpl -> S3ObjectStorage::copyObjectToAnotherObjectStorage.

  • Inline entry: copyFileImpl now reads the bytes with readInlineDataToString (implemented for CAS) and writes them to the destination. S3ObjectStorage throws LOGICAL_ERROR on an empty key.
  • Blob: copyFileImpl gets the offset from the new IMetadataStorage::getObjectPayloadOffset and passes it as object_from_offset. S3ObjectStorage passes it to copyS3File as src_offset, so both the ranged copy and the
    fallback read only the payload.

RESTORE onto CAS is not affected: it writes each part through one CAS transaction.

Changelog category (leave one):

  • Bug Fix (user-visible misbehavior in an official stable release)

Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md):

Fixed BACKUP of a table on a CAS to an S3 and Disk destination on the same endpoint.

Documentation entry for user-facing changes

Adds docs/en/antalya/cas/operations/backup.md describing how BACKUP/RESTORE behave on a content-addressed disk, and links it from the CAS index and roadmap.

CI/CD Options

Exclude tests:

  • Fast test
  • Integration Tests
  • Stateless tests
  • Stateful tests
  • Unit tests
  • Performance tests
  • Aarch64 tests
  • All with ASAN
  • All with TSAN
  • All with MSAN
  • All with UBSAN
  • All with Coverage
  • All Regression
  • Disable CI Cache

Regression jobs to run:

  • Fast suites (mostly <1h)
  • Aggregate Functions (2h)
  • Alter (1.5h)
  • Benchmark (30m)
  • CAS (content-addressed storage; Antalya only)
  • ClickHouse Keeper (1h)
  • Iceberg (2h)
  • LDAP (1h)
  • OAuth (5m)
  • Parquet (1.5h)
  • RBAC (1.5h)
  • SSL Server (1h)
  • S3 (2h)
  • S3 Export (2h)
  • Swarms (30m)
  • Tiered Storage (2h)

Signed-off-by: Konstantin Morozov <just.morozov.k@gmail.com>
@github-actions

github-actions Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Workflow [PR], commit [9544ebf]

Signed-off-by: Konstantin Morozov <just.morozov.k@gmail.com>
Signed-off-by: Konstantin Morozov <just.morozov.k@gmail.com>
Signed-off-by: Konstantin Morozov <just.morozov.k@gmail.com>
@k-morozov
k-morozov marked this pull request as ready for review September 24, 2026 15:50
@k-morozov
k-morozov marked this pull request as draft September 25, 2026 16:26
@k-morozov k-morozov changed the title CAS: backup native copy CAS: S3/Disk backup Sep 25, 2026
@k-morozov k-morozov changed the title CAS: S3/Disk backup BACKUP of CAS tables to S3 on the same endpoint Sep 25, 2026
Signed-off-by: Konstantin Morozov <just.morozov.k@gmail.com>
@filimonov filimonov mentioned this pull request Sep 27, 2026
70 tasks
Signed-off-by: Konstantin Morozov <just.morozov.k@gmail.com>
Signed-off-by: Konstantin Morozov <just.morozov.k@gmail.com>
Signed-off-by: Konstantin Morozov <just.morozov.k@gmail.com>
Signed-off-by: Konstantin Morozov <just.morozov.k@gmail.com>
Signed-off-by: Konstantin Morozov <just.morozov.k@gmail.com>
@k-morozov
k-morozov marked this pull request as ready for review September 29, 2026 06:51

@filimonov filimonov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two things I would change before merge, one is shape and one is a bug.

Shape: the generic API should not learn about envelopes. The PR touches 23 files outside ContentAddressed/ (IDisk, IMetadataStorage, IObjectStorage, S3ObjectStorage, copyS3File, four disk wrappers, ObjectStorageQueue). That is a rebase cost on every release, and the new offset is a defaulted argument that future upstream callers of copyObjectToAnotherObjectStorage / getBlobPath will not know about, so blobs they copy will carry the envelope silently once inline files stop failing first.

The root cause is that DataSourceDescription::operator== and sameKind compare only type, object storage type and endpoint, so a CAS disk looks identical to a plain s3 disk on the same endpoint. Fixing it there keeps everything inside code we already patch:

  • In the DiskObjectStorage constructor (DiskObjectStorage.cpp:124), mark the description for a content-addressed metadata storage (suffix on description, CAS-gated). Then DiskObjectStorage::copyFile falls into IDisk::copyThroughBuffers, and all four backup writers fail sameKind and copy through buffers. Correct bytes, no generic diff. This also covers MOVE out of CAS (CAS-254).
  • copyS3File::performCopy has a real upstream bug: src_offset is ignored when the single-operation CopyObject is chosen. A 3-line fix (never single-op when offset != 0) is independent of CAS and can go upstream on its own.
  • If we want server-side copy for S3(...) destinations, one CAS-gated branch in BackupWriterS3::copyFileFromDisk can call copyS3File with the payload offset as the existing src_offset and a raw readObject fallback reader. No src_object_offset, no getObjectPayloadOffset, no wrapper changes.

Bug: zero-byte blob files. A wide part with an all-empty Array or all-NULL Nullable column has 0-byte arr.bin / n.bin (reproduced with the PR binary). partFileMustStayBlob keeps them blobs, every blob copy is ranged, and calculatePartSize(0) throws LOGICAL_ERROR. BACKUP skips empty files, IDisk::copyFile does not. Needs a size == 0 branch and a test.

Smaller items:

  • getObjectPayloadOffset is only called on the non-CAS branches and always returns 0 there; readInlineDataToString on CAS has no caller (the description says copyFileImpl uses it, it uses getContentAddressedFileCopySource).
  • The blob source is validated in the producer and again in both consumers; src_blob != copy_source.object and payload_size != object.bytes_size cannot be true by construction.
  • No gtest for the ranged copy, although S3ObjectStorageConditionalOpsTest in gtest_writebuffer_s3.cpp already has the mock and a LocalObjectStorage destination (the base IObjectStorage seek path is otherwise untested).
  • GCS: supportsMultiPartCopy is false there, so every blob falls back to the buffered path. Correct, but logged only at TRACE and not tested; worth a line in backend.md.
  • Cost: every blob is now Create + UploadPartCopy + Complete, even a 100-byte primary.idx.

k-morozov and others added 9 commits September 30, 2026 11:46
A column of all-empty `Array` values produces a zero-size `.bin`. Placement
keys on the file name, so `partFileMustStayBlob` keeps it a blob, and a ranged
server-side copy of a blob reaches `calculatePartSize(0)`, which throws.

`BACKUP` never meets such a file when files are deduplicated, but
`IDisk::copyFile` does, so `MOVE PARTITION` out of a CAS disk hits it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`DataSourceDescription::operator==` and `sameKind` compare the storage kind and
the endpoint, so a content-addressed disk looks identical to a plain s3 disk on
the same endpoint. Every server-side copy path then assumes a file is its whole
object starting at byte 0, which is false on a CAS disk: a blob-backed file is a
payload window behind a fixed-size envelope, and an inline file has no object of
its own.

Add `files_are_whole_objects` to `DataSourceDescription` and a predicate
`canUseNativeCopyWith` that requires it from BOTH sides. A precondition holds on
each side separately, so it is a conjunction rather than a comparison: two disks
that both lack the property are not thereby able to use it. `operator==` and
`sameKind` keep their meaning and stay reflexive, so callers that ask about disk
identity are unaffected.

The field is last in the struct because four call sites brace-initialize
`DataSourceDescription` positionally.

`BackupIO_AzureBlobStorage` does not use `sameKind` - it compares
`object_storage_type` directly - so both of its conditions check the capability
explicitly. This closes a reachable configuration rather than only future code: a
read-only CAS mount over Azure is constructible, because `Pool::open` skips the
conditional-write capability probe when the pool is opened read-only, and the
Azure native path copies a whole blob with no range.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Konstantin Morozov <just.morozov.k@gmail.com>
Signed-off-by: Konstantin Morozov <just.morozov.k@gmail.com>
Ten tests for the paths the capability gate touches, most of which no test
exercised before.

Three of them guard non-CAS behaviour, because the gate sits on a path every
object-storage disk takes:

- an ordinary s3 disk against an `encrypted` disk over the same s3. `sameKind`
  ignores `is_encrypted` and `DiskEncrypted` reports its delegate's
  description, so a predicate that replaced `operator==` rather than joining it
  would let the pair through and the cast to `DiskObjectStorage` would throw.
- two ordinary s3 disks keep their server-side copy. The capability defaults to
  false, so any description that forgets to claim it loses the fast path in
  silence.
- `BACKUP TO File(...)` keeps using `fs::copy`. Both sides take their
  description from one function, so a forgotten claim there gives false on both
  sides.

The rest cover CAS: a cache disk over CAS stays out of the server-side copy, a
zero-size file survives a backup when files are not deduplicated, an
incremental backup exercises a non-zero `start_pos`, and a backup written into
a CAS disk now succeeds where it used to be refused.

The mechanism test now also asserts that something went through buffers, so it
cannot pass when a part happens to hold no in-manifest entries.

The unit test pins the predicate's truth table and, separately, that
`operator==` stayed reflexive - the property that breaks if the conjunction is
folded into it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`BACKUP` to an `S3(...)` destination on the pool's own endpoint copies a blob's
payload without its header, and only `UploadPartCopy` can name a byte range. A
store without it is still supported - the blob goes through the ClickHouse
server - but the distinction belongs in the capability table next to the other
store requirements, where an operator choosing a backend will look for it.

`GCS` is the store this applies to today: its current XML API has no
`UploadPartCopy`, so `Client::supportsMultiPartCopy` reports `false`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`test_backup_to_s3_with_empty_array_column` and
`test_incremental_backup_of_a_cas_table` repeat what
`test_backup_to_s3_with_empty_arrays` and `test_incremental_backup_to_s3`
already do, and the existing pair is stronger: the empty-file one asserts that
a zero-size file actually reached the backup writer, rather than only checking
the data after a restore.

Two more cluster round-trips for no new coverage is time every CI run pays.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three fixes from the whole-branch review.

`BackupIO_AzureBlobStorage` went back to checking the capability flags directly
instead of `canUseNativeCopyWith`. The predicate carries `sameKind`, which
compares the description, and the two Azure sides build that string
differently: a disk reports `Endpoint::getServiceEndpoint`, while the backup
reports `ConnectionParams::getConnectionURL`, which for a `connection_string`
disk parses that string and returns the service URL instead. Requiring them
equal would have taken Azure's native copy away from a working configuration
with no error and no log line. The capability is the only dimension this work
argued about, so it is the only one the Azure conditions gained.

`ReadOnlyDiskWrapper` now forwards `getS3StorageClient` and
`tryGetS3StorageClient`. It already forwards `isContentAddressed` and
`getMetadataStorage`, so a CAS disk behind the wrapper reaches the new backup
branch, which then asks the disk for its client and got `NOT_IMPLEMENTED` from
the base class - an exception that escapes the copy rather than falling back to
buffers.

`test_backup_to_file_keeps_fs_copy` asserted that the backup read zero bytes
through a file descriptor. It cannot: `BackupFileInfo` checksums every entry
without a precalculated hash, and `checksums.txt`, `columns.txt`, `count.txt`
and the codec and version files have none. The assertion now compares what was
read against the part's size on disk, which is what "the data did not go
through the server" actually means.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
k-morozov and others added 2 commits September 30, 2026 21:54
`BACKUP TABLE t TO Disk('<cas disk>', ...)` does not work, and the test that
claimed it does was wrong. A backup's own layout mirrors the table's data
directory - `.../data/<db>/<table>/all_1_1_0/<file>` - and
`Cas::isPartFilePath` looks for a part-directory component anywhere in the
path, so a backup file counts as a part file and the autocommit refusal
applies. The test now asserts the refusal, and the documented limitation says a
`CAS` disk cannot be a backup destination rather than claiming the opposite.

`test_backup_to_file_keeps_fs_copy` had its bound the wrong way round. Every
entry pays one checksum pass through `ReadBufferFromFileDescriptor`, so one
pass over the data is what a working `fs::copy` looks like, and the run that
failed - 6832069 bytes read against 6830303 on disk - was the fast path doing
its job. A second pass is what a lost `fs::copy` would cost, so the assertion
is now against one and a half passes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`test_backup_to_s3_with_empty_arrays` counted objects in the backup with
`_size = 0` through the `s3` table function, which cannot report them:
`s3_skip_empty_files` defaults to true, and the listing drops a zero-size
object before the `WHERE` ever sees it. The assertion would have failed even
with the file sitting in the backup.

The query now turns that setting off.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants