You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 410c4eb
Browse filesBrowse the repository at this point in the historyBrowse files
Copy file name to clipboardExpand all lines: mkdocs/docs/api.md
+30Lines changed: 30 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -365,6 +365,34 @@ for buf in tbl.scan().to_arrow_batch_reader():
365
365
print(f"Buffer contains {len(buf)} rows")
366
366
```
367
367
368
+
### Write API modes: `Table` and `Transaction`
369
+
370
+
Every write operation is available through two APIs.
371
+
372
+
The **`Table` API** exposes each operation directly on the table object: `tbl.append(...)`, `tbl.overwrite(...)`, `tbl.delete(...)`, `tbl.dynamic_partition_overwrite(...)` and `tbl.upsert(...)`. Each call opens a transaction, applies the single operation, and commits it as one atomic snapshot. This is the simplest mode and the right default when you only need a single write.
373
+
374
+
The **`Transaction` API** exposes the same operations on a transaction object obtained from `tbl.transaction()`. It batches multiple operations into a single atomic commit: either every operation becomes visible together, or, on failure, none of them do, and readers never observe an intermediate state.
375
+
376
+
```python
377
+
with tbl.transaction() as txn:
378
+
txn.delete(delete_filter="city == 'Paris'")
379
+
txn.append(df_new_cities)
380
+
# both changes commit together when the block exits
381
+
```
382
+
383
+
A transaction can also combine data and metadata changes, for example evolving the schema and writing in the same commit:
If any statement inside the block raises, the whole transaction is discarded and the table is left untouched. The examples in the rest of this section use the `Table` API for brevity; each one has an equivalent method on the transaction object.
395
+
368
396
### Streaming writes from a `RecordBatchReader`
369
397
370
398
`tbl.append()`and `tbl.overwrite()` also accept a `pyarrow.RecordBatchReader` directly, which lets you write datasets that don't fit in memory without materialising them as a `pa.Table` first. PyIceberg consumes the reader once and microbatches it into Parquet files of approximately `write.target-file-size-bytes` (default 512 MiB), keeping memory usage bounded by the target size. All files are committed in a single snapshot.
@@ -386,6 +414,8 @@ df = pa.Table.from_pylist(
386
414
tbl.append(df)
387
415
```
388
416
417
+
### Delete
418
+
389
419
You can delete some of the data from the table by calling `tbl.delete()` with a desired `delete_filter`. This will use the Iceberg metadata to only open up the Parquet files that contain relevant information.
0 commit comments