Skip to content

feat: warn on common GPU env pitfalls before build (#56) - #57

Merged
johnyaku merged 1 commit into
mainfrom
feat/gpu-env-guardrails
Jun 17, 2026
Merged

johnyaku merged 1 commit into
mainfrom
feat/gpu-env-guardrails

Conversation

@johnyaku

@johnyaku johnyaku commented Jun 17, 2026 •

Copy link
Copy Markdown
Contributor

Implements the static guardrails from #56 (items 1, 3, 4). The network-based wheel-triple check (item 2) is tracked separately in #59.

Rebased onto main after #58 merged (clean — no file overlap).

Why

A valid GPU build is a 3-way intersection — (python cpXX) × (framework version) × (CUDA variant cuXXX). Pinning only two leaves the build at the mercy of resolver defaults, and the failure surfaces cryptically ~60s deep in the pip layer (No matching distribution found) rather than up front. This also bit the --base workflow: pointing --base at a CUDA image while leaving pytorch-cuda in the conda deps gives no benefit and still triggers the fragile conda CUDA solve.

What

New static checks in src/absconda/gpu_lint.py (no network — they warn, never block), surfaced at generate/build/publish/validate:

Pitfall Warning
Floating python (>=/unpinned/absent) + a pinned CUDA wheel binary wheels lag new Python releases → pin python=X.Y
Bare torch/jax/… + a CUDA --extra-index-url but no +cuXXX may silently install the CPU wheel (cuda.is_available()==False) → pin +cuXXX
CUDA resolved by conda (pytorch-cuda/cudatoolkit/cuda-*) fragile; message adapts when --base is set — the case that motivated this
Conda pytorch without a CUDA metapackage CPU-only build

Wired into _render_dockerfile (the single choke point that has both the env and the resolved --base) and validate. Docs added to custom-base-images.md (incl. issue #56's note that the conda pytorch channel is not GPU).

Verified against the real cases

Deferred → #59

The network-based (python, torch, cuda) wheel-existence check (item 2 of #56) — querying download.pytorch.org/whl/<cuda>/torch/ to fail fast when no cpXX wheel exists for the pinned interpreter. It's the strongest check but adds a network dependency and needs careful offline/timeout/error-vs-warn UX, so it's tracked in #59. The static floating-python warning here already covers the exact incident in #56.

Testing

  • tests/test_gpu_lint.py — 8 cases (each pitfall + negative/clean envs).
  • Full suite: 109 passed, 2 skipped (docker smoke). ruff check/format clean.

🤖 Generated with Claude Code

A valid GPU build is a 3-way intersection of (python cpXX) x (framework
version) x (cuda variant). Pinning only two leaves the build at the mercy
of resolver defaults, and the failure surfaces cryptically deep in the
pip layer rather than up front. This adds static guardrails that warn
(never block) at generate/build/publish/validate time.

New src/absconda/gpu_lint.py flags:
- Floating python (>=, unpinned, or absent) together with a pinned CUDA
  wheel — binary wheels lag new Python releases, so "newest python" is
  often the combination with no wheel. Recommends pinning python=X.Y.
- A bare binary wheel (torch/jax/cupy/...) alongside a CUDA
  --extra-index-url but no +cuXXX local version — may silently install the
  CPU wheel. Recommends the explicit +cuXXX pin.
- CUDA resolved by conda (pytorch-cuda/cudatoolkit/cuda-*) — fragile, and
  redundant when building on a CUDA --base (the message adapts when --base
  is set, which is the case that motivated this).
- Conda 'pytorch' without a CUDA metapackage — CPU-only build.

Wired into _render_dockerfile (the single choke point that has both the
env and the resolved base) and the validate command. Covers issue #56
guardrails (1), (3), (4); the network-based (python, torch, cuda) wheel
triple check (2) is left as a follow-up. Docs in custom-base-images.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@johnyaku
johnyaku force-pushed the feat/gpu-env-guardrails branch from c8b0219 to c79c2ab Compare June 17, 2026 06:29
@johnyaku
johnyaku merged commit 558cf17 into main Jun 17, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant