Skip to content

GH-41488: [C++][Python] Apply timestamp_parsers as fallback when parsing CSV date and time columns - #50146

Open
pearu wants to merge 7 commits into
apache:mainfrom
pearu:pearu/fix-csv-date-time-parsers
Open

pearu wants to merge 7 commits into
apache:mainfrom
pearu:pearu/fix-csv-date-time-parsers

Conversation

@pearu

@pearu pearu commented Jun 10, 2026 •

Copy link
Copy Markdown
Contributor

Rationale for this change

CSV columns explicitly typed as date32, date64, time32 or time64 can only be parsed from strict ISO-8601 strings: ConvertOptions::timestamp_parsers is consulted only for timestamp columns. Reading e.g. 15-OCT-15 into a date32 column fails even with timestamp_parsers=["%d-%b-%y"], and 7:55:00 (non-zero-padded hour) fails for time32[s]. Users currently work around this by declaring such columns as timestamp, reading, then casting back to the date/time type.

Effect on the issues collected in #41488:

What changes are included in this PR?

  • A new DateTimeWithParsersValueDecoder in csv/converter.cc, used for date32/date64/time32/time64 columns when timestamp_parsers is non-empty. It tries the built-in ISO-8601 parser first (preserving all existing behavior), then each configured parser in order. A timestamp produced by a fallback parser is floored to the day boundary for dates and reduced to the time of day for times, consistent with casting a timestamp to a date or time type. Values carrying a zone offset are rejected, as for zone-less timestamp columns. When no parsers are configured, the pre-existing decoder is used unchanged.
  • Type inference is deliberately unaffected: the Date/Time inference stages now explicitly use options with timestamp_parsers cleared, so inference keeps strict ISO-8601 semantics (otherwise a value with a time-of-day part could be inferred as a date and silently truncated). The existing test_timestamp_parsers Python test pins this behavior.
  • Documentation of the fallback and flooring semantics in ConvertOptions::timestamp_parsers (C++ and Python docstrings) and a new "Date and time parsing" section in the C++ CSV user guide.
  • C-locale name tables for the vendored musl strptime used on Windows, where nl_langinfo() is unavailable. Previously the %a/%A/%b/%B/%h/%p/%c/%r/%x/%X specifiers were compiled out on Windows, so the month-name formats from the original issue reports (%d-%b-%y) could not work there for any column type. The tables match musl's C locale, and name matching is case-insensitive as on glibc/musl/BSD. The fallback path is compiled and verified on Linux via the ARROW_TEST_FALLBACK_LANGINFO hook.

Are these changes tested?

Yes:

  • New C++ tests (Date32Conversion.UserDefinedParsers, Date64Conversion.UserDefinedParsers, Time32Conversion.UserDefinedParsers, Time64Conversion.UserDefinedParsers) covering custom formats, mixed ISO + custom values in one column (backward compatibility of ISO values when parsers are set), pre-epoch flooring with a time-of-day component (distinguishes floor from truncating division), time-of-day extraction from pre-epoch timestamps, zone-offset rejection, and error cases.
  • New Python tests with the reproducers from [C++] Unable to read date64 or date32 in specific format from CSV #28303 and CSV reader cannot parse dates or times #41488, plus an inference-unchanged guard.

Are there any user-facing changes?

Yes: ConvertOptions::timestamp_parsers now also applies, as a fallback after ISO-8601, to columns explicitly typed as date32/date64/time32/time64 (previously such values always errored). No breaking changes: behavior without timestamp_parsers is untouched, ISO values keep parsing when parsers are set, and type inference is unchanged. All language bindings gain the behavior without API changes.

AI usage disclosure

This PR was developed with AI assistance (Claude Code): the decoder, tests and documentation were AI-generated under my direction, then reviewed line-by-line and iterated on by me (design decisions: fallback-after-ISO semantics, silent flooring, inference isolation, and several implementation details adjusted during review). I own and can debug these changes.

🤖 Generated with Claude Code

@pearu

pearu commented Jun 10, 2026

Copy link
Copy Markdown
Contributor Author

Two out-of-scope discoveries made while working on this, recorded here rather than folded into the PR to keep it minimal:

  1. MultipleParsersTimestampValueDecoder::Decode (pre-existing, csv/converter.cc) declares its zone_offset_present flag once outside the parser loop. The built-in parsers happen to write the out-parameter on every call (the strptime parser unconditionally, the ISO-8601 parser resets it to false before scanning), so this is currently harmless — but TimestampParser is a public interface, and a user-implemented parser that only writes the flag when an offset is found could observe a stale value from a previous loop iteration. The new decoder in this PR declares the flag per-iteration; the timestamp decoder could get the same two-line treatment as a MINOR follow-up.

  2. The "Timestamp inference/parsing" section of the C++ CSV user guide (docs/source/cpp/csv.rst) does not mention ConvertOptions::timestamp_parsers at all — custom timestamp parsing was undocumented in the user guide before the date/time subsection added here. A short paragraph there could be a docs follow-up.

@github-actions github-actions Bot added the awaiting review Awaiting review label Jun 10, 2026
@github-actions

Copy link
Copy Markdown

⚠️ GitHub issue #41488 has been automatically assigned in GitHub to PR creator.

@github-actions

Copy link
Copy Markdown

⚠️ GitHub issue #41488 has been automatically assigned in GitHub to PR creator.

@pearu
pearu force-pushed the pearu/fix-csv-date-time-parsers branch from 78ca3cb to e0af29e Compare June 10, 2026 10:08
@pearu

pearu commented Jun 10, 2026 •

Copy link
Copy Markdown
Contributor Author

CI triage of the first run (3 failing jobs, all Windows): all three shared one root cause — the tests used the %d-%b-%y format from the original issue reproducer, but month names (%b) are not supported by the vendored musl strptime used on Windows: cpp/src/arrow/vendored/musl/strptime.c force-undefines HAVE_LANGINFO on _WIN32, which compiles out the %a/%A/%b/%B/%h/%c/%p/%r/%x/%X cases entirely. The feature code is unaffected; numeric-format and time tests passed on Windows.

Fixed (amended) by making numeric formats the primary test coverage and keeping the month-name reproducer guarded to non-Windows (#ifndef _WIN32 in C++, sys.platform != "win32" in Python, following the existing kStrptimeSupportsZone / test_strftime precedents).
UPDATE: sorry for this noise, Claude was too eager to post comments, it is better restrained now.

A third discovery for the list above: this %b limitation is pre-existing and applies equally to timestamp columns with timestamp_parsers on Windows — it just had no CI coverage because the existing timestamp tests only use numeric formats. Could deserve its own issue (either implementing C-locale month names in the vendored strptime, or documenting the limitation in timestamp_parsers docs).

@pearu
pearu force-pushed the pearu/fix-csv-date-time-parsers branch from e0af29e to e75ecec Compare June 10, 2026 13:19
@github-actions

Copy link
Copy Markdown

⚠️ GitHub issue #41488 has been automatically assigned in GitHub to PR creator.

@pearu
pearu marked this pull request as ready for review June 10, 2026 14:13
@pearu
pearu requested review from AlenkaF, raulcd and rok as code owners June 10, 2026 14:13
@pearu

pearu commented Jun 10, 2026

Copy link
Copy Markdown
Contributor Author

The single CI failure (AMD64 Conda C++ AVX2) is unrelated to this PR: Gandiva's TestTime.TestCastTimestampWithTZ fails identically on main since this morning (passing at 4e25461, failing from ca47cd1 on — see e.g. this main run). castTIMESTAMP_utf8 returns 0 for the Canada/Pacific tz name, pointing at tz-database resolution in the CI conda environment — code this PR does not touch.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR extends the CSV reader’s ConvertOptions::timestamp_parsers behavior so that, when columns are explicitly typed as date32/date64/time32/time64, the reader first attempts the existing ISO-8601 parsing and then falls back to the user-provided timestamp parsers (with flooring/extracting semantics consistent with casting). It also keeps type inference strict (ISO-only) to avoid silent truncation.

Changes:

  • Add a new C++ CSV date/time value decoder that falls back to timestamp_parsers after ISO parsing and applies flooring/time-of-day extraction.
  • Ensure CSV type inference for date/time remains ISO-only even when timestamp_parsers are configured.
  • Update C++/Python docs and add C++/Python tests; improve vendored Windows strptime support for C-locale day/month names.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
python/pyarrow/tests/test_csv.py Adds Python coverage for date/time typed columns using timestamp_parsers fallback and inference guard.
python/pyarrow/_csv.pyx Documents the new fallback behavior in Python ConvertOptions docstring.
docs/source/cpp/csv.rst Adds a “Date and time parsing” section documenting fallback + semantics.
cpp/src/arrow/vendored/musl/strptime.c Adds a C-locale nl_langinfo fallback table for Windows/testing to support %b/%B/%p/....
cpp/src/arrow/csv/options.h Documents fallback semantics for timestamp_parsers in C++ API docs.
cpp/src/arrow/csv/inference_internal.h Ensures date/time inference ignores configured timestamp_parsers.
cpp/src/arrow/csv/converter.cc Implements fallback decoder + converter factory changes for date/time types.
cpp/src/arrow/csv/converter_test.cc Adds C++ tests for date/time fallback parsing behavior and edge cases.

Comment thread cpp/src/arrow/csv/converter.cc
Comment thread python/pyarrow/tests/test_csv.py
@github-actions github-actions Bot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Jun 22, 2026
@pearu
pearu force-pushed the pearu/fix-csv-date-time-parsers branch from e75ecec to fcf4f95 Compare June 22, 2026 18:17
@pearu
pearu requested a review from pitrou as a code owner June 22, 2026 18:17
@pitrou

pitrou commented Jun 23, 2026

Copy link
Copy Markdown
Member

@jorisvandenbossche Do you want to take a look at this PR?

@manyi-w

manyi-w commented Jul 2, 2026

Copy link
Copy Markdown

Hi, a small quick question — I noticed that you mentioned multiple issues in the PR description. I was wondering, does this PR fix all of them? Thanks!

@pearu

pearu commented Jul 2, 2026

Copy link
Copy Markdown
Contributor Author

@manyifire Good question — mostly, but not every one in the same way:

So: yes for #28303/#33357 (and #26783/#26224), yes-on-Windows for #31816/#31971, and for #37180 only when timestamp_parsers is set explicitly (not the default).

@pearu
pearu force-pushed the pearu/fix-csv-date-time-parsers branch from fcf4f95 to 808f110 Compare August 10, 2026 13:25
@pearu

pearu commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

Rebased onto current main and re-ran the CSV suite locally — arrow-csv-test passes 278/278, including the four new Date32/Date64/Time32/Time64Conversion.UserDefinedParsers cases.

Status recap for whoever picks this up:

  • Both automated-reviewer comments were addressed in June — std::ranges::transform replaced with a plain index loop, and the %b locale question answered with a reproducer showing the parsing is locale-independent as written.
  • The two failing jobs in the last CI run (AMD64 macOS 15-intel C++ and Python 3) were the same Homebrew infrastructure flake — brew install --formula aws-sdk-cpp hit a /usr/local/Cellar/cmake lock — unrelated to this change. The fresh run on the rebase should clear them.

@pitrou — you routed this to @jorisvandenbossche back in June and it has been quiet since. Is there anything I can do to make this easier to review? If reviewing the CSV converter change together with the vendored strptime C-locale tables is the sticking point, I'm happy to split the strptime part into its own PR.


🤖 Drafted by Claude Code (an AI agent) and reviewed & approved by pearu.

@pearu
pearu force-pushed the pearu/fix-csv-date-time-parsers branch 2 times, most recently from d866653 to 853447b Compare October 1, 2026 21:45
…n parsing CSV date and time columns

CSV columns explicitly typed as date32, date64, time32 or time64 could
only be parsed from strict ISO-8601 strings; ConvertOptions::timestamp_parsers
was consulted only for timestamp columns.

Make the user-defined timestamp parsers act as a fallback for these
column types: the built-in ISO-8601 parser is tried first (preserving
existing behavior), then each configured parser in order. A timestamp
produced by a fallback parser is floored to the day boundary for dates
and reduced to the time of day for times, consistent with casting a
timestamp to a date or time type.

Type inference of date and time columns is deliberately unaffected:
inference keeps using strict ISO-8601 parsing, otherwise a value with a
time-of-day part could be inferred as a date and silently truncated.

Also provide C-locale name tables to the vendored musl strptime used on
Windows, where nl_langinfo() is unavailable: this makes %a/%A/%b/%B/%h/
%p/%c/%r/%x/%X work on Windows (matching musl's C locale), so that the
month-name formats from the original issue reports parse on all
platforms.

Closes apacheGH-28303.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@pearu
pearu force-pushed the pearu/fix-csv-date-time-parsers branch from 853447b to 09ac291 Compare October 2, 2026 09:06
Comment thread cpp/src/arrow/csv/inference_internal.h Outdated
// Date and time inference must not use the user-defined timestamp parsers,
// otherwise a value with a time-of-day (resp. date) part could be inferred
// as a date (resp. time) and be silently truncated.
date_time_options_ = std::make_unique<ConvertOptions>(options);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we avoid copying all of ConvertOptions for every inferred column? Both include_columns and column_types can grow with the number of columns, so this introduces quadratic memory usage. In a local UBSan debug build, reading a one-row CSV with 2,000 projected integer columns used about 69 MB peak memory without custom parsers and 170 MB with [ISO8601].

One small fix would be to add an is_type_inference flag, defaulting to false, to Converter::Make and pass it through to MakeDateTimeConverter. The Date and Time inference cases would call:

return Converter::Make(type, options_, pool,
                       /*is_type_inference=*/true);

Then MakeDateTimeConverter could select the existing decoder directly:

if (is_type_inference || options.timestamp_parsers.empty()) {
  return std::make_shared<ConverterType<T, NumericValueDecoder<T>>>(
      type, options, pool);
}

That would let us remove date_time_options_ entirely while preserving strict date/time inference, custom timestamp inference, and the original options by reference.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, good catch! Applied your suggestion: Converter::Make now takes an is_type_inference flag, and the per-column copy of ConvertOptions is gone.

Peak RSS when reading a one-row CSV with all columns listed in include_columns (release build):

columns no parsers [ISO8601] before [ISO8601] after
1000 67.5 MB 99.9 MB 68.2 MB
2000 76.4 MB 200.5 MB 76.2 MB
4000 95.1 MB 588.0 MB 95.2 MB
8000 133.0 MB 2110.6 MB 133.7 MB

@github-actions github-actions Bot added awaiting changes Awaiting changes and removed awaiting committer review Awaiting committer review labels Oct 6, 2026
// extracted without further conversion
return type.unit();
} else {
return TimeUnit::SECOND;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Using TimeUnit::SECOND here causes the ISO fallback to reject fractional-second timestamps before the date flooring runs. For example:

import pyarrow as pa
import pyarrow.csv as csv

csv.read_csv(
    pa.BufferReader(b"a\n2020-03-15 14:30:00.123\n"),
    convert_options=csv.ConvertOptions(
        column_types={"a": pa.date32()},
        timestamp_parsers=[csv.ISO8601],
    ),
)

This fails with a conversion error, whereas removing .123 succeeds. The same issue affects date64. Under the documented flooring semantics, both inputs should produce 2020-03-15.

Could we accept the fractional component during parsing and discard it when computing the date? Regression tests for both date types, including a pre-epoch value, would help cover this. The fix should also preserve the supported date range; unconditionally switching to nanosecond parsing would narrow it.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, fixed in 4a0fb79. Accepting the fraction inside the parser would mean either parsing in a finer unit, which narrows the date range as you note, or changing ParseTimestampISO8601/TimestampParser. Instead the fractional digits are split off the string before the ISO-8601 parser runs in seconds, so dates keep the full range and 1–9 fractional digits are accepted, with the rest of the value still validated by the parser. Added tests for date32 and date64, including pre-epoch and year-1600 values, all digit counts, and zone offsets after a fraction still being rejected.

Comment thread cpp/src/arrow/csv/converter.cc Outdated
*out = days * kMillisPerDay;
} else {
static_assert(is_time_type<T>::value);
*out = static_cast<value_type>(timestamp - days * ticks_per_day_);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The multiplication days * ticks_per_day_ can overflow even when the parsed timestamp and resulting time of day are both representable.

For example, parsing 1677-09-21 00:12:44 into time64[ns] with timestamp_parsers=[ISO8601] succeeds at the timestamp-parsing step, but UBSan then reports:

signed integer overflow: -106752 * 86400000000000 cannot be represented in type 'int64_t'

Could we compute the time of day using a normalized remainder instead?

int64_t time_of_day = timestamp % ticks_per_day_;
if (time_of_day < 0) {
  time_of_day += ticks_per_day_;
}
*out = static_cast<value_type>(time_of_day);

This preserves the intended behavior for pre-epoch timestamps without the overflowing intermediate multiplication. A regression test using the input above would cover this boundary.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The date branches in this block need bounds checks too. With a custom C++ TimestampParser that reads integer epoch seconds, I reproduced two additional cases:

  • For date32, 185542587187200 seconds corresponds to 2147483648 days. The unchecked cast silently wraps this to -2147483648.
  • For date64, INT64_MAX seconds triggers a UBSan signed-overflow error in days * kMillisPerDay.

Both should return a conversion error because the result is outside the target type’s range.

Could we also check days against the int32_t bounds before narrowing to date32, and use checked multiplication for date64? The normalized-remainder change suggested above addresses the time branch, but these date branches need separate checks.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, applied as suggested: the normalized remainder in 5239526, with 1677-09-21 00:12:44 → time64[ns] as a regression test, and the date32 bounds check plus checked multiplication for date64 in 90ca767. Both now return a conversion error. Since the built-in parsers only read 4-digit years, the range tests use a small test-only parser that reads integer epoch seconds, like your reproduction.

@ianmcook

ianmcook commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

@pearu Thanks for your patience and persistence. I reviewed this and added some comments with assistance from GPT-6 Astra.

if constexpr (is_time_type<T>::value) {
// Parse in the time type's own unit, so that the time of day can be
// extracted without further conversion
return type.unit();

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Using the time type’s unit for the intermediate timestamp unnecessarily restricts time64[ns] to the timestamp[ns] date range.

For example, with input 9999-12-31 07:55:00:

  • Reading directly as time64[ns] fails with a conversion error.
  • Reading as timestamp[s] and casting to time64[ns] succeeds, producing 07:55:00.000000000.

I reproduced this with both ISO8601 and %Y-%m-%d %H:%M:%S, and with year 1600 as well. The resulting time of day is representable; the failure comes from scaling the entire timestamp to nanoseconds before discarding the date.

Could we extract the time at a usable intermediate precision before scaling to the output unit, while preserving fractional precision? Regression tests with dates outside the nanosecond timestamp range would cover this.

This failure occurs during parsing, so replacing the later multiplication with a normalized remainder won’t resolve it.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, fixed in 4e14bcb. The value is still parsed in the time unit first; when that fails, it is parsed again in seconds and only the time of day is scaled to the unit. For the ISO-8601 parser the fractional digits are split off and parsed separately in the time unit, so precision is kept for any date (9999-12-31 07:55:00.123456789 → time64[ns] works). Tests cover years 1600 and 9999 with both ISO-8601 and %Y-%m-%d %H:%M:%S.

pearu and others added 5 commits October 7, 2026 13:59
Pass an is_type_inference flag through Converter::Make to
MakeDateTimeConverter instead of keeping a per-column copy of
ConvertOptions with timestamp_parsers cleared. The copy made memory
usage quadratic in the number of columns when timestamp_parsers was set.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The ISO-8601 fallback parsed dates in seconds, so values with fractional
seconds were rejected before the flooring to the day could run. Discard the
fractional digits before parsing, which keeps the full date range of
parsing in seconds.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
days * ticks_per_day can overflow near the minimum of nanosecond
timestamps even though the time of day is representable. Use a
normalized remainder instead.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Parsing in the time unit limits time64[ns] columns to dates between
1677 and 2262 although the time of day is representable. Retry in
seconds when parsing in the unit fails, and for the ISO-8601 parser
split the fractional seconds off so that they are parsed separately in
the time unit, keeping their precision for any date.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Narrowing the day count to date32 wrapped silently, and scaling it to
milliseconds for date64 could overflow. Both are reachable with a
user-defined timestamp parser, so return a conversion error instead.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions github-actions Bot added awaiting change review Awaiting change review and removed awaiting changes Awaiting changes labels Oct 8, 2026
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Allow ConvertOptions.timestamp_parsers for date types [C++] Unable to read date64 or date32 in specific format from CSV

5 participants