English | 中文
tf-kernel provides optimized CUDA operations for TeleFuser, including fused elementwise operations, quantized
GEMM, SageAttention, and block-sparse attention kernels. It supports SM80, SM90, and SM100 GPU families.
Important
The project does not publish prebuilt tf-kernel wheels or a source distribution to a public package index. Build
and install the extension from source with this directory's Makefile. Direct pip install . and
pip install -e . source builds are intentionally rejected.
- Python 3.10 or newer
- PyTorch 2.11.0
- CUDA Toolkit 12.8 or newer
- CMake 3.26 or newer
- An NVIDIA GPU in the SM80, SM90, or SM100 family
git clone https://github.com/Tele-AI/TeleFuser.git
cd TeleFuser/tf-kernel
make build-auto PYTHON=/path/to/venv/bin/pythonbuild-auto detects the local GPU architecture. The Makefile builds a correctly tagged wheel under dist/ and
installs it into the interpreter selected by PYTHON. This does not install or depend on the TeleFuser Python package.
Local builds use a linux_* platform tag. Only the container build may emit manylinux_2_28 after checking the
wheel's ELF symbol versions against that policy.
Use an explicit architecture target for reproducible builds:
| Target | GPU family |
|---|---|
make build-sm80 |
Ampere and Ada |
make build-sm90 |
Hopper, including H100 |
make build-sm100 |
Blackwell |
make build |
All supported architectures |
For example:
make build-sm90 PYTHON=/path/to/venv/bin/pythonA locally built wheel may be copied to another host or stored in a controlled artifact repository when the source commit, tf-kernel version, PyTorch version, PyTorch CUDA version, C++11 ABI, target SM family, CPU architecture, and Linux/GLIBC baseline are compatible. The wheel validates the runtime facts that can be checked during import.
Architecture-specific builds currently have the same filename for SM80, SM90, and SM100. Keep them in separate artifact paths and install the exact file; do not expose multiple target-SM variants through one simple package index because pip cannot select a wheel from the GPU architecture. For example:
tf-kernel/0.1.0/torch2.11.0-cu128/linux-x86_64/sm90/
Before sharing, run make test-wheel, make test-smoke, and record sha256sum dist/*.whl with the source commit and
test results. Install the selected artifact into an environment that already has the matching PyTorch build:
python -m pip install /path/to/tf_kernel-*.whl --no-deps
python -m pip checkDo not manually relabel a local linux_* wheel as manylinux; rebuild it on the intended deployment baseline.
See the full installation guide for the artifact manifest and target-host verification procedure.
MAX_JOBS controls concurrent build jobs. TF_KERNEL_COMPILE_THREADS controls NVCC threads within each job:
make build-auto \
PYTHON=/path/to/venv/bin/python \
MAX_JOBS=16 \
TF_KERNEL_COMPILE_THREADS=4Higher values can reduce build time on a sufficiently provisioned host, but also increase CPU and memory pressure.
For a resource-constrained build, start with MAX_JOBS=2 TF_KERNEL_COMPILE_THREADS=1.
Run the smoke test with the same interpreter passed to Make:
/path/to/venv/bin/python - <<'PY'
from pathlib import Path
import torch
import tf_kernel
print("tf-kernel:", tf_kernel.__version__)
print("PyTorch:", torch.__version__)
print("GPU:", torch.cuda.get_device_name())
print("extension:", Path(tf_kernel.common_ops.__file__).resolve())
x = torch.randn(8, 1024, device="cuda", dtype=torch.float16)
weight = torch.ones(1024, device="cuda", dtype=torch.float16)
assert torch.isfinite(tf_kernel.rmsnorm(x, weight)).all()
print("RMSNorm smoke test: OK")
PYFP4 operators require SM100 or newer, so FP4 is unavailable on Ampere and Hopper. On H100, sageattn()
and TeleFuser SAGE_ATTN_2_8_8_SM90 use the validated SM90 FP8 implementation when the SM90 wheel is installed.
At import time, the wheel verifies the PyTorch public version, PyTorch CUDA version, C++11 ABI, and target GPU family
recorded during its build. A process must expose GPUs from only one architecture family. SageAttention v2 dispatch is
currently enabled for SM80, SM86, SM89, SM90, SM120, and SM121; other compute capabilities fail explicitly instead of
selecting an unvalidated backend.
make test-cpu PYTHON=/path/to/venv/bin/python
make test-smoke PYTHON=/path/to/venv/bin/python
make test PYTHON=/path/to/venv/bin/python # bounded GPU suite
make test-full PYTHON=/path/to/venv/bin/python # exhaustive matrix
make test-wheel PYTHON=/path/to/venv/bin/python
make format-check PYTHON=/path/to/venv/bin/python
make docs PYTHON=/path/to/venv/bin/pythonGPU targets install the wheel into an isolated temporary directory before collection, so tests cannot accidentally
import tf_kernel from the source checkout or another environment.
See CONTRIBUTING.md for development and wheel distribution policy. See the full installation and usage guide for compatibility, API examples, and troubleshooting.