Skip to content

Commit 410c4eb

Browse files
committed
Docs: document the Transaction write API and add a Delete subsection (#1008)
1 parent c86eb8e commit 410c4eb

1 file changed

Lines changed: 30 additions & 0 deletions

File tree

‎mkdocs/docs/api.md‎

Lines changed: 30 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -365,6 +365,34 @@ for buf in tbl.scan().to_arrow_batch_reader():
365365
print(f"Buffer contains {len(buf)} rows")
366366
```
367367

368+
### Write API modes: `Table` and `Transaction`
369+
370+
Every write operation is available through two APIs.
371+
372+
The **`Table` API** exposes each operation directly on the table object: `tbl.append(...)`, `tbl.overwrite(...)`, `tbl.delete(...)`, `tbl.dynamic_partition_overwrite(...)` and `tbl.upsert(...)`. Each call opens a transaction, applies the single operation, and commits it as one atomic snapshot. This is the simplest mode and the right default when you only need a single write.
373+
374+
The **`Transaction` API** exposes the same operations on a transaction object obtained from `tbl.transaction()`. It batches multiple operations into a single atomic commit: either every operation becomes visible together, or, on failure, none of them do, and readers never observe an intermediate state.
375+
376+
```python
377+
with tbl.transaction() as txn:
378+
txn.delete(delete_filter="city == 'Paris'")
379+
txn.append(df_new_cities)
380+
# both changes commit together when the block exits
381+
```
382+
383+
A transaction can also combine data and metadata changes, for example evolving the schema and writing in the same commit:
384+
385+
```python
386+
from pyiceberg.types import LongType
387+
388+
with tbl.transaction() as txn:
389+
with txn.update_schema() as update_schema:
390+
update_schema.add_column("population", LongType())
391+
txn.append(df_with_population)
392+
```
393+
394+
If any statement inside the block raises, the whole transaction is discarded and the table is left untouched. The examples in the rest of this section use the `Table` API for brevity; each one has an equivalent method on the transaction object.
395+
368396
### Streaming writes from a `RecordBatchReader`
369397

370398
`tbl.append()` and `tbl.overwrite()` also accept a `pyarrow.RecordBatchReader` directly, which lets you write datasets that don't fit in memory without materialising them as a `pa.Table` first. PyIceberg consumes the reader once and microbatches it into Parquet files of approximately `write.target-file-size-bytes` (default 512 MiB), keeping memory usage bounded by the target size. All files are committed in a single snapshot.
@@ -386,6 +414,8 @@ df = pa.Table.from_pylist(
386414
tbl.append(df)
387415
```
388416

417+
### Delete
418+
389419
You can delete some of the data from the table by calling `tbl.delete()` with a desired `delete_filter`. This will use the Iceberg metadata to only open up the Parquet files that contain relevant information.
390420

391421
```python

0 commit comments

Comments
 (0)