Repository navigation
improve Spark from_utc_timestamp compatibility - #25979
Open
nat-openai wants to merge 1 commit into
Open
nat-openai wants to merge 1 commit into
nat-openai wants to merge 1 commit into
Conversation
Author
|
@sunchao, if this seems reasonable, can you enable the CI workflows? |
The timezone spellings accepted by Spark are different than those accepted by Arrow's parser. As a result, `from_utc_timestamp` in comet isn't safe to use, since it would fail on things like Java short IDs, prefixed offsets with second precision, and SystemV-style IDs. Additionally, Arrow will accept offsets beyond the Spark limit of 18 hours. This change improves compatibility by parsing timezone IDs with the rules implemented by Spark. As part of doing this, I added SparkFromUtcTimestampExpr so that computed timezones evaluated for rows with non-null timestamps, avoiding failing when a timezone is invalid but wouldn't be required anyway. (Full compatibility with spark behaviour seems quite complex and I haven't attempted it here.) It also adds more documentation recording remaining differences: - timezone validation and evaluation order can differ from Spark - chrono accepts a narrower calendar range than Spark implemented range - chrono-tz doesn't currently perform DST transitions past 2099 This does not achieve full compatibility with Spark due to those issues, but it makes the region of incompatibility a lot smaller.
nat-openai
force-pushed
the
spark-timezone-support
branch
from
October 9, 2026 17:49
3244992 to
7594f11
Compare
ajsquared
approved these changes
Oct 9, 2026
ajsquared
left a comment
There was a problem hiding this comment.
Reviewed the current changes and feedback; no independently confirmed unresolved P1 remains.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue does this PR close?
This addresses the major incompatibility described in apache/datafusion-comet#4654 for
CometFromUTCTimestamp. It does not close that issue.It's a follow-up to #19879.
Rationale for this change
The timezone spellings accepted by Spark are different than those accepted by Arrow's parser. As a result,
from_utc_timestampin comet isn't safe to use, since it would fail on things like Java short IDs, prefixed offsets with second precision, and SystemV-style IDs. Additionally, Arrow will accept offsets beyond the Spark limit of 18 hours.What changes are included in this PR?
This change improves compatibility by parsing timezone IDs with the rules implemented by Spark. As part of doing this, I added SparkFromUtcTimestampExpr so that computed timezones evaluated for rows with non-null timestamps, avoiding failing when a timezone is invalid but wouldn't be required anyway. (Full compatibility with spark behaviour seems quite complex and I haven't attempted it here.)
It also adds more documentation recording remaining differences:
This does not achieve full compatibility with Spark due to those issues, but it makes the region of incompatibility a lot smaller.
If this approach isn't the one that folks would prefer, please let me know if there are different avenues you'd rather I explore!
Finally, it's worth mentioning that this change was LLM-assisted.
What is the testing strategy for this PR?
Added unit tests in
datafusion/spark/src/function/datetime/timezone.rscovering accepted spelling, the 18 hour boundary, and SystemV custom behaviour around DST transitions and exception.Added SQL logic tests in
datafusion/sqllogictest/test_files/spark/datetime/from_utc_timestamp.sltcovering the accepted spellings.Are there any user-facing changes?
Yes. The existing Spark
from_utc_timestampUDF now accepts additional Spark timezone spellings. It also rejects offsets outside Spark's 18-hour limit.This won't affect most users since they'd need to have opted into the existing incompatible implementation.