Repository navigation
feat: warn on common GPU env pitfalls before build (#56) - #57
Merged
Merged
Conversation
A valid GPU build is a 3-way intersection of (python cpXX) x (framework version) x (cuda variant). Pinning only two leaves the build at the mercy of resolver defaults, and the failure surfaces cryptically deep in the pip layer rather than up front. This adds static guardrails that warn (never block) at generate/build/publish/validate time. New src/absconda/gpu_lint.py flags: - Floating python (>=, unpinned, or absent) together with a pinned CUDA wheel — binary wheels lag new Python releases, so "newest python" is often the combination with no wheel. Recommends pinning python=X.Y. - A bare binary wheel (torch/jax/cupy/...) alongside a CUDA --extra-index-url but no +cuXXX local version — may silently install the CPU wheel. Recommends the explicit +cuXXX pin. - CUDA resolved by conda (pytorch-cuda/cudatoolkit/cuda-*) — fragile, and redundant when building on a CUDA --base (the message adapts when --base is set, which is the case that motivated this). - Conda 'pytorch' without a CUDA metapackage — CPU-only build. Wired into _render_dockerfile (the single choke point that has both the env and the resolved base) and the validate command. Covers issue #56 guardrails (1), (3), (4); the network-based (python, torch, cuda) wheel triple check (2) is left as a follow-up. Docs in custom-base-images.md. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
johnyaku
force-pushed
the
feat/gpu-env-guardrails
branch
from
June 17, 2026 06:29
c8b0219 to
c79c2ab
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements the static guardrails from #56 (items 1, 3, 4). The network-based wheel-triple check (item 2) is tracked separately in #59.
Why
A valid GPU build is a 3-way intersection — (python
cpXX) × (framework version) × (CUDA variantcuXXX). Pinning only two leaves the build at the mercy of resolver defaults, and the failure surfaces cryptically ~60s deep in the pip layer (No matching distribution found) rather than up front. This also bit the--baseworkflow: pointing--baseat a CUDA image while leavingpytorch-cudain the conda deps gives no benefit and still triggers the fragile conda CUDA solve.What
New static checks in
src/absconda/gpu_lint.py(no network — they warn, never block), surfaced atgenerate/build/publish/validate:python(>=/unpinned/absent) + a pinned CUDA wheelpython=X.Ytorch/jax/… + a CUDA--extra-index-urlbut no+cuXXXcuda.is_available()==False) → pin+cuXXXpytorch-cuda/cudatoolkit/cuda-*)--baseis set — the case that motivated thispytorchwithout a CUDA metapackageWired into
_render_dockerfile(the single choke point that has both the env and the resolved--base) andvalidate. Docs added tocustom-base-images.md(incl. issue #56's note that the condapytorchchannel is not GPU).Verified against the real cases
flash-scope(python>=3.11+torch==2.2.2+cu118) → floating-python warning.torch+ cu118 index → silent-CPU warning.cell2location(condapytorch-cuda+--base) → base-aware conda-CUDA warning.cell2location(pinnedpython=3.10,torch==2.2.2+cu118) → no false positives.Deferred → #59
The network-based (python, torch, cuda) wheel-existence check (item 2 of #56) — querying
download.pytorch.org/whl/<cuda>/torch/to fail fast when nocpXXwheel exists for the pinned interpreter. It's the strongest check but adds a network dependency and needs careful offline/timeout/error-vs-warn UX, so it's tracked in #59. The static floating-python warning here already covers the exact incident in #56.Testing
tests/test_gpu_lint.py— 8 cases (each pitfall + negative/clean envs).ruff check/formatclean.🤖 Generated with Claude Code