Skip to content

Pull requests: NVlabs/cuda-oxide

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

feat(cuda-device): add f16x2 and bf16x2 unpacking to f32
#476 opened Jul 25, 2026 by honeyspoon Contributor Loading…
feat(cuda-device): add f32 and f64 warp reduce utilities
#475 opened Jul 25, 2026 by honeyspoon Contributor Loading…
feat(cuda-device): derive the m16n8k16 accumulator fragment layout
#474 opened Jul 25, 2026 by honeyspoon Contributor Loading…
fix(mir): keep small runtime-indexed arrays in SSA bug Something isn't working build-related Build scripts, toolchain discovery, or compilation setup codegen Device code-generation pipeline (Rust MIR to IR to PTX) enhancement New feature or request IR-lowering Lowering between dialects (dialect-mir to LLVM dialect) on-hold Paused: valid but deferred until a prerequisite decision or dependency resolves perf Performance of generated code or of the compiler itself
#398 opened Jul 13, 2026 by niklebedenko Contributor Draft
fix(mir-lower): keep scalar float math on the LLVM-to-PTX path enhancement New feature or request intrinsics Device intrinsics and libdevice math mappings IR-lowering Lowering between dialects (dialect-mir to LLVM dialect) on-hold Paused: valid but deferred until a prerequisite decision or dependency resolves perf Performance of generated code or of the compiler itself
#391 opened Jul 13, 2026 by niklebedenko Contributor Loading…
6 tasks done
feat: fp and cuda-graph support
#346 opened Jul 6, 2026 by devillove084 Loading…
1 of 2 tasks
ProTip! Type g p on any issue or pull request to go back to the pull request listing page.