Description
What would you like to see added?
Ray Data currently supports reading an Iceberg table via ray.data.read_iceberg(...), including reading a specific snapshot with snapshot_id. However, it does not appear to expose Iceberg's incremental read capability for reading data between two snapshots.
It would be useful for Ray Data to support incremental reads for Iceberg tables, for example by allowing users to read appended rows between a start snapshot and an end snapshot.
PyIceberg has implemented incremental append scan support in apache/iceberg-python#3512:
apache/iceberg-python#3512
That PR adds Table.incremental_append_scan(...), which can scan rows appended by append snapshots within a snapshot range.
Why is this needed?
Many data pipelines only need to process newly appended Iceberg data since the last processed snapshot. Without native support in Ray Data, users need to manually inspect Iceberg snapshot history/manifests or read full snapshots and compute diffs externally, which is inefficient and error-prone.
Native support would make it easier to build incremental ETL and ML data processing pipelines on top of Ray Data and Iceberg.
Possible API
One possible direction is to extend ray.data.read_iceberg(...) with incremental scan options, such as:
ray.data.read_iceberg(
table_identifier="db.table",
catalog_kwargs={...},
start_snapshot_id=..., # exclusive or inclusive depending on Iceberg semantics
end_snapshot_id=...,
)
Alternatively, Ray Data could expose a separate API such as ray.data.read_iceberg_incremental(...).
Additional context
Current Ray Data Iceberg support passes snapshot_id to PyIceberg Table.scan(), which enables reading a specific snapshot but not reading the delta between snapshots. Since PyIceberg now has incremental append scan support, Ray Data could potentially build on top of that functionality.
cc @ray-project/data
Use case
No response
Description
What would you like to see added?
Ray Data currently supports reading an Iceberg table via
ray.data.read_iceberg(...), including reading a specific snapshot withsnapshot_id. However, it does not appear to expose Iceberg's incremental read capability for reading data between two snapshots.It would be useful for Ray Data to support incremental reads for Iceberg tables, for example by allowing users to read appended rows between a start snapshot and an end snapshot.
PyIceberg has implemented incremental append scan support in apache/iceberg-python#3512:
apache/iceberg-python#3512
That PR adds
Table.incremental_append_scan(...), which can scan rows appended by append snapshots within a snapshot range.Why is this needed?
Many data pipelines only need to process newly appended Iceberg data since the last processed snapshot. Without native support in Ray Data, users need to manually inspect Iceberg snapshot history/manifests or read full snapshots and compute diffs externally, which is inefficient and error-prone.
Native support would make it easier to build incremental ETL and ML data processing pipelines on top of Ray Data and Iceberg.
Possible API
One possible direction is to extend
ray.data.read_iceberg(...)with incremental scan options, such as:Use case
No response