Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
32 commits
Select commit Hold shift + click to select a range
5a1ed93
add test to reproduce
k-morozov Sep 22, 2026
2fcdc9e
fix empty blob_key, fix blob's payload offset
k-morozov Sep 24, 2026
8fe883f
update doc
k-morozov Sep 24, 2026
b130c47
update test
k-morozov Sep 24, 2026
6d6674f
support S3 disk
k-morozov Sep 25, 2026
11e2cf1
update tests
k-morozov Sep 28, 2026
1fb8ca8
use seekable buffer
k-morozov Sep 28, 2026
38d8d7c
separate cas source
k-morozov Sep 28, 2026
a6e026f
general checks
k-morozov Sep 28, 2026
24c9f87
use visit
k-morozov Sep 28, 2026
f2a4858
Merge branch 'antalya-26.6' into cas/backup-native-copy-bug
k-morozov Sep 30, 2026
0dedc56
test: MOVE PARTITION out of a CAS disk with an empty-array column
k-morozov Sep 30, 2026
dbaa406
Gate native copy on a positive whole-object capability
k-morozov Sep 30, 2026
b1050f4
apply comments
k-morozov Sep 30, 2026
50f37d5
fix bugs
k-morozov Sep 30, 2026
a3e96df
Cover the whole-object capability and the CAS S3 copy branch with tests
k-morozov Sep 30, 2026
cbc9f62
Document the ranged same-store copy as a bucket capability
k-morozov Sep 30, 2026
869d795
Drop two tests that duplicate existing coverage
k-morozov Sep 30, 2026
18d0837
Keep the Azure gate on the capability alone, and forward the S3 client
k-morozov Sep 30, 2026
d912c65
Fix the three integration failures CI found
k-morozov Sep 30, 2026
9544ebf
Let the zero-size assertion see the object it looks for
k-morozov Sep 30, 2026
9490050
update doc
k-morozov Oct 1, 2026
3cb053d
change check order
k-morozov Oct 1, 2026
17fd068
add events to test
k-morozov Oct 1, 2026
68af947
update tests
k-morozov Oct 1, 2026
6853a19
update tests
k-morozov Oct 1, 2026
7ce84b3
update tests
k-morozov Oct 1, 2026
fe961fc
update tests
k-morozov Oct 1, 2026
312527a
update doc
k-morozov Oct 1, 2026
866ddab
update doc
k-morozov Oct 1, 2026
6ed1e11
Merge branch 'antalya-26.6' into cas/backup-native-copy-bug
k-morozov Oct 1, 2026
90dbe76
fix CI
k-morozov Oct 1, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions docs/en/antalya/cas/bucket-requirements.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,7 @@ these conditions is refused rather than trusted.
| Unconditional complete-object publication | `Backend::publishBlob` | An absent or condemned content-addressed body is replaced atomically; native stores may use multipart |
| Native same-store copy when `cas_staging_backend = s3` | `IObjectStorage::copyObject` with `ObjectStorageCopyMode::NativeOnly` | The first absent staged publication may copy its complete object without a client-side fallback |
| Exact-token delete | `Backend::deleteExact` | GC must delete only the incarnation it condemned, never a replacement |
| Ranged same-store copy, for backups only | `UploadPartCopy` through `copyS3File` | A `BACKUP` to an `S3(...)` destination on the pool's own endpoint copies a blob's payload without its header. Optional, but only when the lack is recognized: `GCS` lacks it by its current XML API, so `Client::supportsMultiPartCopy` reports `false` there and the copy goes through the ClickHouse server, which is slower but correct. A store that instead refuses `UploadPartCopy` with an error other than `AccessDenied` fails the `BACKUP`; set `s3_allow_multipart_copy = 0` there |
| Ranged `GET` | `Backend::get` / `Backend::getStream` with a `Range` | Opening one column file of a part costs one bounded read, not a whole-object fetch |
| `LIST` with a resumable cursor | `Backend::list` | GC discovery and the orphan-manifest sweep page through the pool without a separate index |
| No versioning / no delete markers | probed by `runCapabilityProbe`; `created_delete_marker` on `DeleteOutcome` | A delete marker over a live key would break exact-token semantics — GC would archive instead of reclaim |
Expand Down
1 change: 1 addition & 0 deletions docs/en/antalya/cas/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,4 +94,5 @@ per disk, so adopting it never requires migrating an existing deployment.
| [Architecture overview](/antalya/cas/architecture/) | The object model, the Git analogy, and the safety invariants |
| [Correctness](/antalya/cas/architecture/correctness) | How the design was verified: TLA+ models, counterexamples, soak methodology |
| [Design history](/antalya/cas/architecture/design-history) | What earlier designs were tried and rejected, and why |
| [Backup](/antalya/cas/operations/backup) | How `BACKUP` and `RESTORE` behave on a content-addressed disk |
| [Roadmap](/antalya/cas/roadmap) | What is shipped, planned, and deliberately not pursued |
150 changes: 150 additions & 0 deletions docs/en/antalya/cas/operations/backup.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,150 @@
---
description: 'How BACKUP and RESTORE work for a table on a content-addressed disk: what holds the data during a backup, when the copy runs inside S3, and what is not supported yet.'
sidebar_label: 'Backup'
sidebar_position: 5
slug: /antalya/cas/operations/backup
title: 'CAS Operations — Backup'
doc_type: 'guide'
---

# Operations — backup {#backup}

Ordinary `BACKUP` and `RESTORE` work for a table on a content-addressed (`CAS`) disk. This page
covers what happens during one, how it differs from a plain disk, and what the limits are.

The `CAS`-native backup model — `snapshot` / `mirror` / `fetch` / `restore` — is designed but not
implemented and is not wired into the SQL surface. See the [roadmap](/antalya/cas/roadmap#backups).

## What is supported {#supported}

```sql
BACKUP TABLE t TO S3('https://bucket.s3.amazonaws.com/backups/b1', 'key', 'secret');
RESTORE TABLE t AS t_restored FROM S3('https://bucket.s3.amazonaws.com/backups/b1', 'key', 'secret');
```

The destination can be anything: `S3`, `Disk`, `File`, or an archive. A backup can be restored onto a
disk of any type, because it holds the table's files rather than the pool's objects.

**An `Atomic` database is required.** That has been the default since 20.x. On the deprecated
`Ordinary` engine the backup fails with `SUPPORT_IS_DISABLED`: that path pins files with temporary
hard links, which object storage does not have.

## What holds the data during a backup {#holding}

On a plain disk a backup pins files against deletion with a hard link. `CAS` uses a different
mechanism — pointer holding:

- the backup holds a `shared_ptr` to the table and to each part;
- the outdated-part cleanup skips those parts;
- while a part is alive so is its [ref](/antalya/cas/architecture/manifests-and-refs#ref-table) — the
name under which the part is registered in its namespace's ref table, and through which it points
at its manifest;
- while the ref is alive, garbage collection sees the manifest and its blobs as reachable.

This mechanism lives in the process's memory and **does not survive a server restart**. An
interrupted backup leaves nothing behind in the pool, but it also stops protecting the data once the
process is gone. For durable pinning there is `FREEZE`, which publishes a real ref.

## How the bytes move {#copy-path}

Which path runs depends on the destination.

**A destination outside the pool** — the common case: another bucket, a local disk, an archive. Files
are read through the `CAS` read path and written to the destination. Pool deduplication is lost:
what was one blob shared by several replicas becomes ordinary files in the backup.

**An `S3(...)` destination on the same `S3` endpoint as the pool** — the copy of a blob then runs
inside the `S3` store itself, without sending the bytes through ClickHouse. The server issues one `UploadPartCopy` per upload part, so a payload larger than one part takes several.

```sql
BACKUP TABLE t TO S3('http://s3.example.com/bucket/backups/b1', 'key', 'secret');
```

Here the pool is under `http://s3.example.com/bucket/pool/`.

The files of a part fall into two categories:

| Category | Example | How it is copied |
|---|---|---|
| Blob | `data.bin`, marks, `primary.idx` | an `UploadPartCopy` naming a byte range — only the payload moves, without the blob's internal header |
| Inside the manifest | `checksums.txt`, `count.txt`, `columns.txt` | read through the `CAS` read path and written through ClickHouse's buffers |

A blob object is `[header][payload]`, so a file never starts at the beginning of its object. Every
copy of a blob therefore names the payload range. A copy of the whole object would put the header
into the backup, and a later restore would read wrong data.

If a byte range cannot be copied inside `S3`, `CAS` does not fall back to copying the whole object —
the file is read and written through ClickHouse instead. That is slower, but correct. ClickHouse takes
that path on its own in three cases: the query sets `s3_allow_multipart_copy = 0`; the store is
recognized as `GCS`, whose current XML API has no `UploadPartCopy`; or `UploadPartCopy` is refused with
`AccessDenied`.

A store that refuses `UploadPartCopy` with any other error fails the `BACKUP` instead, because
ClickHouse cannot tell a missing capability from a transient failure. On such a store, set
`s3_allow_multipart_copy = 0` on the `BACKUP` query.

| Turn off the copy inside `S3` | Turn off the range copy |
|---|---|
| `SETTINGS allow_s3_native_copy = 0` in the `BACKUP` query | the query setting `s3_allow_multipart_copy = 0` |

```sql
BACKUP TABLE t TO S3(...) SETTINGS allow_s3_native_copy = 0;
```

**A `Disk(...)` destination** never copies inside `S3`, even when the disk is an `s3` or `s3_plain`
disk on the same endpoint as the pool:

```sql
BACKUP TABLE t TO Disk('backups_s3', 'b1');
```

Every file is read through the `CAS` read path and written through ClickHouse's buffers. The
`allow_s3_native_copy` and `s3_allow_multipart_copy` settings have no effect on this path.

## Restore {#restore}

Each part is materialized in **one disk transaction** and published as one manifest and one ref. A
partially restored part can never appear in the pool: either the whole part is published or nothing
is.

Restore onto a `CAS` disk never copies objects inside `S3`, even when the backup is on the same
endpoint as the pool. Each file of the part is read from the backup and written through the `CAS`
write path, because only that path can build the manifest and the blobs. A restore onto a plain disk
works as usual and can copy inside `S3`.

Restored data is packed afresh — on a `CAS` disk it gets new blobs and new refs. Deduplication
against data already in the pool works as usual: identical content hashes to the same blob and is
not written twice.

## `FREEZE` is not a backup {#freeze}

`ALTER TABLE ... FREEZE` works on `CAS` and publishes parts into a separate shadow namespace, which
is a garbage-collection root in its own right. `DROP PARTITION` removes the live refs and leaves the
snapshot alone; `SYSTEM UNFREEZE` removes only the shadow refs.

It is still not a snapshot of a table: there is no SQL metadata, no single commit marker, no
portable object with a listing and a restore API, and its lifetime is tied to a manual `UNFREEZE`.
It is a useful building block, not a replacement for `BACKUP`.

## Limitations {#limitations}

- The `CAS`-native backup model (`snapshot` / `mirror` / `fetch`) is not implemented.
- The `Ordinary` database engine is not supported.
- Pool deduplication is lost in the backup: its size follows the logical files, not the unique blobs.
- Pointer holding does not survive a server restart.
- The copy inside the store works only for an `S3(...)` destination on an `S3` or `S3`-compatible
store, and only when the destination is on the same endpoint as the pool. Other destinations,
including `Disk(...)`, get the copy through ClickHouse's buffers.
- A blob is copied inside `S3` only with multipart copy (`UploadPartCopy`), because only it can name a
byte range. With `s3_allow_multipart_copy = 0`, on a store recognized as `GCS`, and when
`UploadPartCopy` is refused with `AccessDenied`, every blob goes through ClickHouse's buffers
instead. A store that refuses `UploadPartCopy` with another error fails the `BACKUP`; set
`s3_allow_multipart_copy = 0` there.
- Restore onto a `CAS` disk always writes through ClickHouse, see [restore](#restore).
- A disk-level write of a part file onto a `CAS` disk outside of a part transaction is rejected with
`NOT_IMPLEMENTED` and the message `Autocommit writes are not supported for content part files`: a
`CAS` disk publishes a part's manifest and ref at commit, so it has nowhere to put a single
autocommitted file.
- A `CAS` disk cannot be a backup destination: `BACKUP TABLE t TO Disk('<cas disk>', 'b1')` is rejected
the same way. A backup's own layout mirrors the table's data directory, so its files sit under a part
directory as well and count as part files.
4 changes: 4 additions & 0 deletions docs/en/antalya/cas/roadmap.md
Original file line number Diff line number Diff line change
Expand Up @@ -91,6 +91,10 @@ positioning.

## Backups {#backups}

Ordinary `BACKUP` and `RESTORE` already work for a table on a `CAS` disk — see
[backup](/antalya/cas/operations/backup) for how they behave and what the limits are. What follows is
about the `CAS`-native model, which is a different thing.

A `snapshot` / `mirror` / `fetch` / `restore` design is **approved but not implemented**. The
model is deliberately git-shaped: `snapshot` is instant and free (like `git tag` — it references
existing manifests, copies nothing); `mirror` is a continuous pull from a production pool into a
Expand Down
28 changes: 24 additions & 4 deletions src/Backups/BackupIO_AzureBlobStorage.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,15 @@ BackupReaderAzureBlobStorage::BackupReaderAzureBlobStorage(
const WriteSettings & write_settings_,
const ContextPtr & context_)
: BackupReaderDefault(read_settings_, write_settings_, getLogger("BackupReaderAzureBlobStorage"))
, data_source_description{DataSourceType::ObjectStorage, ObjectStorageType::Azure, MetadataStorageType::None, connection_params_.getConnectionURL(), false, false, ""}
, data_source_description{
.type = DataSourceType::ObjectStorage,
.object_storage_type = ObjectStorageType::Azure,
.metadata_type = MetadataStorageType::None,
.description = connection_params_.getConnectionURL(),
.is_encrypted = false,
.is_cached = false,
.zookeeper_name = "",
.files_are_whole_objects = true}
, connection_params(connection_params_)
, blob_path(blob_path_)
{
Expand Down Expand Up @@ -87,7 +95,9 @@ void BackupReaderAzureBlobStorage::copyFileToDisk(const String & path_in_backup,
auto destination_data_source_description = destination_disk->getDataSourceDescription();
LOG_TRACE(log, "Source description {}, destination description {}", data_source_description.description, destination_data_source_description.description);
if (destination_data_source_description.object_storage_type == ObjectStorageType::Azure
&& destination_data_source_description.is_encrypted == encrypted_in_backup)
&& destination_data_source_description.is_encrypted == encrypted_in_backup
&& destination_data_source_description.files_are_whole_objects
&& data_source_description.files_are_whole_objects)
{
LOG_TRACE(log, "Copying {} from AzureBlobStorage to disk {}", path_in_backup, destination_disk->getName());
auto write_blob_function = [&](const Strings & dst_blob_path, WriteMode mode, const std::optional<ObjectAttributes> &) -> size_t
Expand Down Expand Up @@ -133,7 +143,15 @@ BackupWriterAzureBlobStorage::BackupWriterAzureBlobStorage(
const ContextPtr & context_,
bool attempt_to_create_container)
: BackupWriterDefault(read_settings_, write_settings_, getLogger("BackupWriterAzureBlobStorage"))
, data_source_description{DataSourceType::ObjectStorage, ObjectStorageType::Azure, MetadataStorageType::None, connection_params_.getConnectionURL(), false, false, ""}
, data_source_description{
.type = DataSourceType::ObjectStorage,
.object_storage_type = ObjectStorageType::Azure,
.metadata_type = MetadataStorageType::None,
.description = connection_params_.getConnectionURL(),
.is_encrypted = false,
.is_cached = false,
.zookeeper_name = "",
.files_are_whole_objects = true}
, connection_params(connection_params_)
, blob_path(blob_path_)
{
Expand Down Expand Up @@ -165,7 +183,9 @@ void BackupWriterAzureBlobStorage::copyFileFromDisk(
auto source_data_source_description = src_disk->getDataSourceDescription();
LOG_TRACE(log, "Source description {}, destination description {}", source_data_source_description.description, data_source_description.description);
if (source_data_source_description.object_storage_type == ObjectStorageType::Azure
&& source_data_source_description.is_encrypted == copy_encrypted)
&& source_data_source_description.is_encrypted == copy_encrypted
&& source_data_source_description.files_are_whole_objects
&& data_source_description.files_are_whole_objects)
{
/// getBlobPath() can return more than 2 elements if the file is stored as multiple objects in AzureBlobStorage container.
/// In this case we can't use the native copy.
Expand Down
Loading
Loading