TruncateTransform splits multi-byte UTF-8 characters, producing invalid string bounds in manifests
Impact
TruncateTransform.Transformer for strings does a byte-based cut (v[:min(len(v), t.Width)]), so truncating a lower bound to 16 bytes can split a multi-byte UTF-8 character mid-sequence. The resulting manifest contains an invalid-UTF-8 lower_bounds entry. Strict readers (verified: Cloudflare R2 SQL / DataFusion-based engine, error cannot decode the Iceberg manifest body using its metadata) then refuse to decode the manifest, and because later snapshots inherit the manifest, one bad bound makes the whole table unreadable.
Repro
Any string column whose 16-byte truncation boundary falls inside a multi-byte character. Real case from production data (Indonesian news, content column starting with POSKOTA.CO.ID… where … is U+2026 e2 80 a6):
- full min value:
POSKOTA.CO.ID… (504f534b4f54412e434f2e494420e280a6...)
- manifest
lower_bounds written: 504f534b4f54412e434f2e494420e280 (16 bytes, trailing partial e2 80)
The writer path is statsAggregator.MinAsBytes (table/internal/utils.go) applying TruncateTransform{Width: truncLen} (default metrics mode truncate(16)), whose string case in transforms.go is:
case string:
return v[:min(len(v), t.Width)]
Note the upper-bound path (TruncateUpperBoundText) is character-aware, but the lower-bound path via TruncateTransform is not.
Expected
Lower-bound truncation must cut on a UTF-8 character boundary (like the upper-bound path does), so manifests never contain invalid string bounds.
Environment
iceberg-go v0.6.0, table v2, writer path table.Append with default table properties.
TruncateTransform splits multi-byte UTF-8 characters, producing invalid string bounds in manifests
Impact
TruncateTransform.Transformerfor strings does a byte-based cut (v[:min(len(v), t.Width)]), so truncating a lower bound to 16 bytes can split a multi-byte UTF-8 character mid-sequence. The resulting manifest contains an invalid-UTF-8lower_boundsentry. Strict readers (verified: Cloudflare R2 SQL / DataFusion-based engine, errorcannot decode the Iceberg manifest body using its metadata) then refuse to decode the manifest, and because later snapshots inherit the manifest, one bad bound makes the whole table unreadable.Repro
Any string column whose 16-byte truncation boundary falls inside a multi-byte character. Real case from production data (Indonesian news, content column starting with
POSKOTA.CO.ID…where…is U+2026e2 80 a6):POSKOTA.CO.ID…(504f534b4f54412e434f2e494420e280a6...)lower_boundswritten:504f534b4f54412e434f2e494420e280(16 bytes, trailing partiale2 80)The writer path is
statsAggregator.MinAsBytes(table/internal/utils.go) applyingTruncateTransform{Width: truncLen}(default metrics modetruncate(16)), whose string case intransforms.gois:Note the upper-bound path (
TruncateUpperBoundText) is character-aware, but the lower-bound path viaTruncateTransformis not.Expected
Lower-bound truncation must cut on a UTF-8 character boundary (like the upper-bound path does), so manifests never contain invalid string bounds.
Environment
iceberg-go v0.6.0, table v2, writer path
table.Appendwith default table properties.