Skip to content

Latest commit

 

History

History
139 lines (104 loc) · 5.56 KB

File metadata and controls

139 lines (104 loc) · 5.56 KB

tf-kernel

English | 中文

License Python CUDA PyTorch

tf-kernel provides optimized CUDA operations for TeleFuser, including fused elementwise operations, quantized GEMM, SageAttention, and block-sparse attention kernels. It supports SM80, SM90, and SM100 GPU families.

Important

The project does not publish prebuilt tf-kernel wheels or a source distribution to a public package index. Build and install the extension from source with this directory's Makefile. Direct pip install . and pip install -e . source builds are intentionally rejected.

Requirements

  • Python 3.10 or newer
  • PyTorch 2.11.0
  • CUDA Toolkit 12.8 or newer
  • CMake 3.26 or newer
  • An NVIDIA GPU in the SM80, SM90, or SM100 family

Build and Install

git clone https://github.com/Tele-AI/TeleFuser.git
cd TeleFuser/tf-kernel
make build-auto PYTHON=/path/to/venv/bin/python

build-auto detects the local GPU architecture. The Makefile builds a correctly tagged wheel under dist/ and installs it into the interpreter selected by PYTHON. This does not install or depend on the TeleFuser Python package. Local builds use a linux_* platform tag. Only the container build may emit manylinux_2_28 after checking the wheel's ELF symbol versions against that policy.

Use an explicit architecture target for reproducible builds:

Target GPU family
make build-sm80 Ampere and Ada
make build-sm90 Hopper, including H100
make build-sm100 Blackwell
make build All supported architectures

For example:

make build-sm90 PYTHON=/path/to/venv/bin/python

Distribute a Built Wheel

A locally built wheel may be copied to another host or stored in a controlled artifact repository when the source commit, tf-kernel version, PyTorch version, PyTorch CUDA version, C++11 ABI, target SM family, CPU architecture, and Linux/GLIBC baseline are compatible. The wheel validates the runtime facts that can be checked during import.

Architecture-specific builds currently have the same filename for SM80, SM90, and SM100. Keep them in separate artifact paths and install the exact file; do not expose multiple target-SM variants through one simple package index because pip cannot select a wheel from the GPU architecture. For example:

tf-kernel/0.1.0/torch2.11.0-cu128/linux-x86_64/sm90/

Before sharing, run make test-wheel, make test-smoke, and record sha256sum dist/*.whl with the source commit and test results. Install the selected artifact into an environment that already has the matching PyTorch build:

python -m pip install /path/to/tf_kernel-*.whl --no-deps
python -m pip check

Do not manually relabel a local linux_* wheel as manylinux; rebuild it on the intended deployment baseline. See the full installation guide for the artifact manifest and target-host verification procedure.

Parallel Compilation

MAX_JOBS controls concurrent build jobs. TF_KERNEL_COMPILE_THREADS controls NVCC threads within each job:

make build-auto \
  PYTHON=/path/to/venv/bin/python \
  MAX_JOBS=16 \
  TF_KERNEL_COMPILE_THREADS=4

Higher values can reduce build time on a sufficiently provisioned host, but also increase CPU and memory pressure. For a resource-constrained build, start with MAX_JOBS=2 TF_KERNEL_COMPILE_THREADS=1.

Verify

Run the smoke test with the same interpreter passed to Make:

/path/to/venv/bin/python - <<'PY'
from pathlib import Path

import torch
import tf_kernel

print("tf-kernel:", tf_kernel.__version__)
print("PyTorch:", torch.__version__)
print("GPU:", torch.cuda.get_device_name())
print("extension:", Path(tf_kernel.common_ops.__file__).resolve())

x = torch.randn(8, 1024, device="cuda", dtype=torch.float16)
weight = torch.ones(1024, device="cuda", dtype=torch.float16)
assert torch.isfinite(tf_kernel.rmsnorm(x, weight)).all()
print("RMSNorm smoke test: OK")
PY

FP4 operators require SM100 or newer, so FP4 is unavailable on Ampere and Hopper. On H100, sageattn() and TeleFuser SAGE_ATTN_2_8_8_SM90 use the validated SM90 FP8 implementation when the SM90 wheel is installed. At import time, the wheel verifies the PyTorch public version, PyTorch CUDA version, C++11 ABI, and target GPU family recorded during its build. A process must expose GPUs from only one architecture family. SageAttention v2 dispatch is currently enabled for SM80, SM86, SM89, SM90, SM120, and SM121; other compute capabilities fail explicitly instead of selecting an unvalidated backend.

Development

make test-cpu PYTHON=/path/to/venv/bin/python
make test-smoke PYTHON=/path/to/venv/bin/python
make test PYTHON=/path/to/venv/bin/python       # bounded GPU suite
make test-full PYTHON=/path/to/venv/bin/python  # exhaustive matrix
make test-wheel PYTHON=/path/to/venv/bin/python
make format-check PYTHON=/path/to/venv/bin/python
make docs PYTHON=/path/to/venv/bin/python

GPU targets install the wheel into an isolated temporary directory before collection, so tests cannot accidentally import tf_kernel from the source checkout or another environment.

See CONTRIBUTING.md for development and wheel distribution policy. See the full installation and usage guide for compatibility, API examples, and troubleshooting.