-
Notifications
You must be signed in to change notification settings - Fork 223
Pull requests: NVlabs/cuda-oxide
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
feat(cuda-device): add f16x2 and bf16x2 unpacking to f32
#476
opened Jul 25, 2026 by
honeyspoon
Contributor
Loading…
feat(cuda-device): add f32 and f64 warp reduce utilities
#475
opened Jul 25, 2026 by
honeyspoon
Contributor
Loading…
feat(cuda-device): derive the m16n8k16 accumulator fragment layout
#474
opened Jul 25, 2026 by
honeyspoon
Contributor
Loading…
feat: implement core::intrinsics::exact_div (unblocks slice::as_chunks)
#473
opened Jul 25, 2026 by
honeyspoon
Contributor
Loading…
fix(mir): keep small runtime-indexed arrays in SSA
bug
Something isn't working
build-related
Build scripts, toolchain discovery, or compilation setup
codegen
Device code-generation pipeline (Rust MIR to IR to PTX)
enhancement
New feature or request
IR-lowering
Lowering between dialects (dialect-mir to LLVM dialect)
on-hold
Paused: valid but deferred until a prerequisite decision or dependency resolves
perf
Performance of generated code or of the compiler itself
#398
opened Jul 13, 2026 by
niklebedenko
Contributor
•
Draft
fix(mir-lower): keep scalar float math on the LLVM-to-PTX path
enhancement
New feature or request
intrinsics
Device intrinsics and libdevice math mappings
IR-lowering
Lowering between dialects (dialect-mir to LLVM dialect)
on-hold
Paused: valid but deferred until a prerequisite decision or dependency resolves
perf
Performance of generated code or of the compiler itself
#391
opened Jul 13, 2026 by
niklebedenko
Contributor
Loading…
6 tasks done
ProTip!
Type g p on any issue or pull request to go back to the pull request listing page.