Skip to content

[Data][Docs] Document disk-based shuffle in Data internals - #66488

Merged
owenowenisme merged 5 commits into
ray-project:masterfrom
owenowenisme:owenowenisme/Add-Doc-about-disk-shuffle
Sep 27, 2026
Merged

owenowenisme merged 5 commits into
ray-project:masterfrom
owenowenisme:owenowenisme/Add-Doc-about-disk-shuffle

Conversation

@owenowenisme

Copy link
Copy Markdown
Member

Description

Add a disk-based shuffle subsection under shuffle v2: what it is, the per-node file-server actor serving shards over Arrow Flight, when to use it, and how to enable it.

Related issues

Link related issues: "Fixes #1234", "Closes #1234", or "Related to #1234".

Additional information

Optional: Add implementation details, API changes, usage examples, screenshots, etc.

Add a disk-based shuffle subsection under shuffle v2: what it is, the
per-node file-server actor serving shards over Arrow Flight, when to
use it, and how to enable it.

Signed-off-by: You-Cheng Lin <c-youcheng.lin@anyscale.com>
Signed-off-by: You-Cheng Lin <mses010108@gmail.com>
…e APIs

Add a tip to Dataset.repartition (keys), Dataset.join, and
Dataset.groupby docstrings pointing to the disk-based shuffle section
for datasets larger than aggregate object-store memory.

Signed-off-by: You-Cheng Lin <mses010108@gmail.com>
Signed-off-by: You-Cheng Lin <mses010108@gmail.com>
@owenowenisme
owenowenisme requested review from a team as code owners September 25, 2026 06:02
@owenowenisme owenowenisme added docs An issue or change related to documentation data Ray Data-related issues docs-go RtD-only checks for docs-only changes. Doesn't run full library doc-test suites. labels Sep 25, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces documentation and docstring tips for the new disk-based shuffle feature in Ray Data, explaining its mechanics and configuration. The review feedback recommends updating the API references in the docstrings of repartition, join, and groupby to use the fully qualified, copy-pasteable ray.data.DataContext.get_current().use_disk_based_hash_shuffle = True instead of the shorthand DataContext.use_disk_based_hash_shuffle = True.

Comment thread python/ray/data/dataset.py
Comment thread python/ray/data/dataset.py
Comment thread python/ray/data/dataset.py
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Signed-off-by: You-Cheng Lin <106612301+owenowenisme@users.noreply.github.com>
@richardliaw richardliaw added the go add ONLY when ready to merge, run all tests label Sep 25, 2026
@owenowenisme
owenowenisme enabled auto-merge (squash) September 27, 2026 03:50
@github-actions
github-actions Bot disabled auto-merge September 27, 2026 03:50
@owenowenisme owenowenisme removed the docs-go RtD-only checks for docs-only changes. Doesn't run full library doc-test suites. label Sep 27, 2026
@owenowenisme
owenowenisme force-pushed the owenowenisme/Add-Doc-about-disk-shuffle branch from 5b67749 to 57cf0e6 Compare September 27, 2026 04:27
@owenowenisme
owenowenisme merged commit 1cf93ad into ray-project:master Sep 27, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

data Ray Data-related issues docs An issue or change related to documentation go add ONLY when ready to merge, run all tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ray fails to serialize self-reference objects

2 participants