Reported by @mbutrovich in #19.
Unpartitioned tables: NULLs are written into required columns
The unpartitioned path skips project_with_partition (table/mod.rs), and IcebergWriteExec passes the input's own schema as the sink schema to execute_input_stream (write.rs).
DataFusion only runs check_not_null_constraints for columns that are non-nullable in the sink schema and nullable in the input. Because both schemas are the input's here, the check never runs.
Repro: inserting SELECT * FROM source into an unpartitioned table with a required id: int column, from a MemTable whose nullable id holds [1, NULL], succeeds and reads back 0, 1. The NULL is written as 0.
Partitioned tables: nullable sources are rejected at plan time
The partitioned path goes the other way. project_with_partition rejects any source column that is nullable where the table column is required, even when the source holds no nulls. DataFusion's own sinks accept that case and check at runtime.
Possible fix
Passing the table's Arrow schema (plus the partition column) as the sink schema to execute_input_stream would likely fix both: nulls would be caught at runtime on either path, and the plan-time rejection of nullable sources could be relaxed.
Reported by @mbutrovich in #19.
Unpartitioned tables: NULLs are written into required columns
The unpartitioned path skips
project_with_partition(table/mod.rs), andIcebergWriteExecpasses the input's own schema as the sink schema toexecute_input_stream(write.rs).DataFusion only runs
check_not_null_constraintsfor columns that are non-nullable in the sink schema and nullable in the input. Because both schemas are the input's here, the check never runs.Repro: inserting
SELECT * FROM sourceinto an unpartitioned table with a requiredid: intcolumn, from aMemTablewhose nullableidholds[1, NULL], succeeds and reads back0, 1. The NULL is written as0.Partitioned tables: nullable sources are rejected at plan time
The partitioned path goes the other way.
project_with_partitionrejects any source column that is nullable where the table column is required, even when the source holds no nulls. DataFusion's own sinks accept that case and check at runtime.Possible fix
Passing the table's Arrow schema (plus the partition column) as the sink schema to
execute_input_streamwould likely fix both: nulls would be caught at runtime on either path, and the plan-time rejection of nullable sources could be relaxed.