fix: accept NOT NULL input columns in partitioned INSERT - #19
NoahKusaba wants to merge 2 commits into
Conversation
mbutrovich
left a comment
There was a problem hiding this comment.
Thanks @NoahKusaba! The fix looks right to me, and I confirmed that test_insert_not_null_source_into_partitioned_table fails with main's project.rs and passes with this change.
While testing the partitioned and unpartitioned paths side by side, I found a gap on the unpartitioned side that this PR doesn't cause but that sits next to its motivation. Could you open an issue for it? The unpartitioned path skips project_with_partition (table/mod.rs), and IcebergWriteExec passes the input's own schema as the sink schema to execute_input_stream (write.rs). DataFusion only runs its runtime null check (check_not_null_constraints) for columns that are non-nullable in the sink schema and nullable in the input. Because both schemas are the input's here, the check never runs. At the head commit, inserting SELECT * FROM source into an unpartitioned table with a required id: int column, from a MemTable whose nullable id holds [1, NULL], succeeds and reads back 0, 1. The NULL is written as 0. The partitioned path goes the other way and rejects a nullable source at plan time even when it holds no nulls, which DataFusion's own sinks accept and check at runtime. Passing the table's Arrow schema (plus the partition column) as the sink schema would probably fix both, so the issue could cover both.
| fn test_schema_validation_nested_nullability() { | ||
| let child = |nullable| Field::new("x", DataType::Int32, nullable); | ||
| let input = |nullable| { | ||
| input_of(vec![ | ||
| Field::new("id", DataType::Int32, false), | ||
| Field::new( | ||
| "s", | ||
| DataType::Struct(Fields::from(vec![child(nullable)])), | ||
| false, | ||
| ), | ||
| ]) | ||
| }; | ||
| let table = |x: NestedField| { | ||
| table_partitioned_by_id(vec![ | ||
| NestedField::required(1, "id", Type::Primitive(PrimitiveType::Int)), | ||
| NestedField::required( | ||
| 2, | ||
| "s", | ||
| Type::Struct(StructType::new(vec![Arc::new(x)])), | ||
| ), | ||
| ]) | ||
| }; | ||
| let int = Type::Primitive(PrimitiveType::Int); | ||
|
|
||
| // The same rule applies inside a struct. | ||
| let optional = table(NestedField::optional(3, "x", int.clone())); | ||
| assert!(project_with_partition(input(false), &optional).is_ok()); | ||
| assert!(project_with_partition(input(true), &optional).is_ok()); | ||
|
|
||
| let required = table(NestedField::required(3, "x", int)); | ||
| assert!(project_with_partition(input(false), &required).is_ok()); | ||
| let err = project_with_partition(input(true), &required) | ||
| .unwrap_err() | ||
| .to_string(); | ||
| assert!(err.contains(INCOMPATIBLE), "{err}"); | ||
| } |
There was a problem hiding this comment.
The new doc says the rule holds "at any nesting depth", and this test covers a struct field. Could we add the same four cases for a list element and a map value? Lists and maps go through their own arms of Arrow's DataType::contains, and the map arm also compares the keys_sorted flag, so a struct test doesn't cover them.
There was a problem hiding this comment.
Thanks, good call. Adding the list and map cases turned up a bug already on main: iceberg-rust's strip_metadata_from_schema fails on any list or map column with "Field stack underflow in list", so a partitioned INSERT with such a column errors out before the schemas are compared. MetadataStripVisitor only pushes onto its field stack in before_field, which the visitor doesn't call for list elements or map keys and values. Filed as apache/iceberg-rust#3297.
The fix and a regression test are ready. I'll open that PR once my pending iceberg-rust PRs are merged (apache/iceberg-rust#3286, apache/iceberg-rust#2904), since reviews there are bandwidth-limited.
In abee042 the four cases live in a shared assert_nested_nullability helper used by the struct test, and the list and map tests are staged but commented out until the fix lands. They pass against 665c64e with the fix applied, including the map case with keys_sorted = false (what Iceberg maps convert to), so enabling them is just uncommenting.
Two list/map limitations of contains I noticed while writing them, neither new in this PR (the old == had both):
- Field names are compared, so a list whose element is named
item(Arrow's default, used byDataType::new_list) won't match Iceberg'selement. - A source map with
keys_sorted = trueis rejected, since Iceberg maps convert to unsorted Arrow maps.
If you think either is worth handling, I can fold them into #22 or open a separate issue.
Move the four nullability cases into assert_nested_nullability and run them for a struct field. Add the same cases for a list element and a map value, commented out until iceberg-rust's strip_metadata_from_schema supports lists and maps; today it errors on them before the schemas are compared.
|
@mbutrovich Opened #22 for the unpartitioned gap, with your repro. It covers both sides:
I'd like to keep this PR to the narrower fix, accepting NOT NULL sources into optional columns, and leave the plan-time rejection as it is. Once #22 passes the table's schema as the sink schema, that rejection can be dropped in favour of DataFusion's runtime check, and both paths will behave the same. |
Which issue does this PR close?
What changes are included in this PR?
An
INSERTinto a partitioned table fails when the source has aNOT NULLcolumn where the table's column is optional:Every value of a non-nullable column is valid in an optional one, so the write is safe. The same
INSERTinto an unpartitioned table already succeeds, becauseproject_with_partitionreturns before this check when the spec is unpartitioned, so today the result depends on whether the table is partitioned.project_with_partitioncompared the input and table schemas with==. It now uses Arrow'sSchema::contains, which allows the input to be narrower in nullability and is otherwise as strict as before:The public doc of
project_with_partitionnow states this contract. The error message changes from "does not match" to "is not compatible with", since the schemas no longer need to be equal.The new unit tests need a partitioned table with a given schema. The three existing schema-validation tests each built one inline with the same ~50 lines, so they now share a
table_partitioned_by_idhelper, with their imports moved to the top of thetestsmodule. Their assertions are unchanged apart from the new error message.Are these changes tested?
test_insert_not_null_source_into_partitioned_table(integration): inserts from aNOT NULLMemTableinto a table partitioned on an optional column, and reads the rows back. It fails onmainwith the error above.test_schema_validation_nullability: non-nullable and nullable input into an optional column are both accepted; into a required column, non-nullable is accepted and nullable is rejected.test_schema_validation_struct_nullability: the same four cases for a field inside a struct.cargo fmt --all -- --check,cargo clippy --workspace --locked --all-targets -- -D warningsandcargo test --workspace --lockedall pass locally.AI Disclosure
main.🤖 Generated with Claude Code