Skip to content

GH-51760: [C++][Python] Emit decimal points for integral floats - #51767

Open
gitedmond wants to merge 1 commit into
apache:mainfrom
gitedmond:fix-float-string-decimal-point
Open

gitedmond wants to merge 1 commit into
apache:mainfrom
gitedmond:fix-float-string-decimal-point

Conversation

@gitedmond

@gitedmond gitedmond commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Rationale for this change

Fixes #51760.

Casting an integral floating-point value to a string currently produces 20 rather than 20.0. The CSV writer uses that cast, so a first fragment containing only whole-number floats is inferred as int64; scanning a later fragment with fractional values then fails.

What changes are included in this PR?

Enable EMIT_TRAILING_DECIMAL_POINT and EMIT_TRAILING_ZERO_AFTER_POINT in the default floating-point formatter, as suggested in the issue.

Extend the existing float16/float32/float64 formatter and cast tests, update affected scalar and pretty-print expectations, and add a CSV dataset regression covering float32/float64 inputs with threaded and serial scans. The regression also checks that integer columns retain integer output and inference.

Are these changes tested?

Built Arrow C++ and PyArrow from the patched source. Removing only the two flags and rebuilding reproduced the exact reported int64 conversion error and failed all four regression cases; restoring them made the original example and the regression pass.

  • C++: 19 formatting tests, all 108 cast tests, 168 scalar tests, 193 miscellaneous tests (including pretty-printing), and all 280 CSV tests passed.
  • Python: 17 CSV dataset tests and 132 CSV tests passed; 10 tests in each selection skipped because cloudpickle and optional compression support were unavailable.
  • clang-format 18.1.8, flake8, and git diff --check passed.

Validation used a Release C++ build with compute, CSV, dataset, filesystem, IPC, and JSON enabled; RE2, utf8proc, and optional compression libraries were disabled.

Are there any user-facing changes?

Integral floats in ordinary decimal notation now include .0 in string casts, CSV output, scalar strings, and array display, including -0.0. Fractional values, scientific notation (for example, 1e+10), infinity, NaN, and integer formatting keep their existing spelling.

Was AI used for this PR?

AI assisted with the implementation, tests, validation, and description. The production change follows the formatter flags proposed by the issue author.

PR code and description written by:

  • Human
  • AI

Reviewed before submission by:

  • Human
  • AI
  • Not reviewed

@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

⚠️ GitHub issue #51760 has been automatically assigned in GitHub to PR creator.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[C++][Python] Make float-to-string cast write decimal point on integral values

1 participant