Skip to content

improve Spark from_utc_timestamp compatibility - #25979

Open
nat-openai wants to merge 1 commit into
apache:mainfrom
nat-openai:spark-timezone-support
Open

nat-openai wants to merge 1 commit into
apache:mainfrom
nat-openai:spark-timezone-support

Conversation

@nat-openai

Copy link
Copy Markdown

Which issue does this PR close?

This addresses the major incompatibility described in apache/datafusion-comet#4654 for CometFromUTCTimestamp. It does not close that issue.

It's a follow-up to #19879.

Rationale for this change

The timezone spellings accepted by Spark are different than those accepted by Arrow's parser. As a result, from_utc_timestamp in comet isn't safe to use, since it would fail on things like Java short IDs, prefixed offsets with second precision, and SystemV-style IDs. Additionally, Arrow will accept offsets beyond the Spark limit of 18 hours.

What changes are included in this PR?

This change improves compatibility by parsing timezone IDs with the rules implemented by Spark. As part of doing this, I added SparkFromUtcTimestampExpr so that computed timezones evaluated for rows with non-null timestamps, avoiding failing when a timezone is invalid but wouldn't be required anyway. (Full compatibility with spark behaviour seems quite complex and I haven't attempted it here.)

It also adds more documentation recording remaining differences:

  • timezone validation and evaluation order can differ from Spark
  • chrono accepts a narrower calendar range than Spark implemented range
  • chrono-tz doesn't currently perform DST transitions past 2099

This does not achieve full compatibility with Spark due to those issues, but it makes the region of incompatibility a lot smaller.

If this approach isn't the one that folks would prefer, please let me know if there are different avenues you'd rather I explore!

Finally, it's worth mentioning that this change was LLM-assisted.

What is the testing strategy for this PR?

Added unit tests in datafusion/spark/src/function/datetime/timezone.rs covering accepted spelling, the 18 hour boundary, and SystemV custom behaviour around DST transitions and exception.

Added SQL logic tests in datafusion/sqllogictest/test_files/spark/datetime/from_utc_timestamp.slt covering the accepted spellings.

Are there any user-facing changes?

Yes. The existing Spark from_utc_timestamp UDF now accepts additional Spark timezone spellings. It also rejects offsets outside Spark's 18-hour limit.

This won't affect most users since they'd need to have opted into the existing incompatible implementation.

@github-actions github-actions Bot added sqllogictest SQL Logic Tests (.slt) functions Changes to functions implementation spark labels Oct 2, 2026
@nat-openai

Copy link
Copy Markdown
Author

@sunchao, if this seems reasonable, can you enable the CI workflows?
Also cc @andygrove, since this is tangentially relevant to apache/datafusion-comet#4654.

The timezone spellings accepted by Spark are different than those
accepted by Arrow's parser. As a result, `from_utc_timestamp` in comet
isn't safe to use, since it would fail on things like Java short IDs,
prefixed offsets with second precision, and SystemV-style IDs.
Additionally, Arrow will accept offsets beyond the Spark limit of 18
hours.

This change improves compatibility by parsing timezone IDs with the
rules implemented by Spark. As part of doing this, I added
SparkFromUtcTimestampExpr so that computed timezones evaluated for rows
with non-null timestamps, avoiding failing when a timezone is invalid
but wouldn't be required anyway. (Full compatibility with spark
behaviour seems quite complex and I haven't attempted it here.)

It also adds more documentation recording remaining differences:

- timezone validation and evaluation order can differ from Spark
- chrono accepts a narrower calendar range than Spark implemented range
- chrono-tz doesn't currently perform DST transitions past 2099

This does not achieve full compatibility with Spark due to those issues,
but it makes the region of incompatibility a lot smaller.
@nat-openai
nat-openai force-pushed the spark-timezone-support branch from 3244992 to 7594f11 Compare October 9, 2026 17:49

@ajsquared ajsquared left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the current changes and feedback; no independently confirmed unresolved P1 remains.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

functions Changes to functions implementation spark sqllogictest SQL Logic Tests (.slt)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants