Thank you for your interest in contributing to tf-kernel! This document provides guidelines and instructions for contributing to this project.
- Code of Conduct
- Getting Started
- Development Setup
- How to Contribute
- Coding Standards
- Commit Message Guidelines
- Pull Request Process
This project adheres to the Contributor Covenant Code of Conduct. By participating, you are expected to uphold this code.
- Fork the repository on GitHub
- Clone your fork locally:
git clone https://github.com/YOUR_USERNAME/TeleFuser.git cd TeleFuser/tf-kernel - Set up the upstream remote:
git remote add upstream https://github.com/Tele-AI/TeleFuser.git
- Python >= 3.10
- CMake >= 3.26
- CUDA Toolkit 12.8+
- PyTorch == 2.11.0
Install development, testing, documentation, and linting tools without invoking a source installation through pip, then compile tf-kernel with Make:
# Install development tools
python -m pip install pytest pytest-cov sphinx sphinx-rtd-theme \
sphinx-autodoc-typehints pre-commit isort ruff clang-format
# Install pre-commit hooks
pre-commit install
# Build and install the project (requires CUDA)
make build-auto PYTHON=/path/to/venv/bin/python
# Run tests
make testIf you only need specific dependencies:
# For running tests
python -m pip install pytest pytest-cov
# For building documentation
python -m pip install sphinx sphinx-rtd-theme sphinx-autodoc-typehints
# For code linting and formatting
python -m pip install pre-commit isort ruff clang-format
# Install pre-commit hooks (after installing lint dependencies)
pre-commit installDo not use pip install ., pip install -e ., or pip install -e ".[dev]" for local tf-kernel source. Direct source
installation fails during CMake configuration. Compilation and installation must go through a make build-* target,
which installs the generated wheel.
| Group | Description | Packages |
|---|---|---|
dev |
All development dependencies | Includes test, docs, and lint |
test |
Testing dependencies | pytest, pytest-cov |
docs |
Documentation dependencies | sphinx, sphinx-rtd-theme, sphinx-autodoc-typehints |
lint |
Code formatting and linting | pre-commit, black, isort, ruff |
Note: clang-format for C++/CUDA formatting should be installed via system package manager:
# Ubuntu/Debian
sudo apt-get install clang-format
# macOS
brew install clang-formatDo not publish prebuilt tf-kernel wheels or a source distribution to PyPI or another public package index. Direct source installation is intentionally unsupported, so publishing an sdist would make installers attempt a build path that the project rejects. Do not add package-index installation instructions or a TeleFuser optional dependency.
Users may build a wheel on an explicitly provisioned CUDA/NVCC host and distribute it directly within a compatible, controlled environment. Before sharing an artifact:
make build-sm90 PYTHON=/path/to/venv/bin/python # Select the required architecture.
make test-wheel PYTHON=/path/to/venv/bin/python
make test-smoke PYTHON=/path/to/venv/bin/python
sha256sum dist/*.whlRecord the source commit, package version, Python version, PyTorch version, PyTorch CUDA version, C++11 ABI, target SM
family, CPU architecture, Linux/GLIBC baseline, wheel SHA-256, and test results with the artifact. Keep SM80, SM90,
and SM100 wheels in separate artifact paths because architecture-specific builds currently produce the same wheel
filename.
Do not place multiple target-SM variants of one release in a single simple package index: pip cannot select a wheel
from the target GPU architecture. Install a shared wheel by exact path or URL with --no-deps, then run
python -m pip check and the installation smoke test on the target host.
Before creating a bug report, please check the existing issues to avoid duplicates.
When filing a bug report, please include:
- System information: OS, CUDA version, GPU architecture
- Python version and PyTorch version
- Steps to reproduce the issue
- Expected behavior vs actual behavior
- Error messages or stack traces
- Minimal code example that reproduces the issue
Enhancement suggestions are tracked as GitHub issues. When creating an enhancement suggestion:
- Use a clear and descriptive title
- Provide a detailed description of the proposed feature
- Explain why this enhancement would be useful
- List some examples of how the feature would be used
To add a new CUDA kernel:
- Implement the kernel in
csrc/<category>/your_kernel.cu - Declare the interface in
include/tf_kernel_ops.h - Register with PyTorch in
csrc/common_extension.cc:m.def("your_kernel(Tensor input, Tensor! output) -> ()"); m.impl("your_kernel", torch::kCUDA, &your_kernel);
- Update CMakeLists.txt: Add source file to the appropriate
SOURCESlist - Create Python wrapper in
tf_kernel/<category>.py - Export in
tf_kernel/__init__.py - Add tests in
tests/test_your_kernel.py - Add benchmarks in
benchmark/(if applicable)
See AGENTS.md for more detailed technical information.
- Use clang-format with the provided
.clang-formatconfig (Google style based) - 2-space indentation
- 120 column limit
- Left pointer alignment (
int* ptrnotint *ptr)
# Format C++/CUDA files
make format- isort: Import sorting
black: Code formatting(Disabled)- ruff: Linting (F821 rule, F401 disabled)
# Format Python files
make formatAll commits are checked by pre-commit hooks. You can run them manually:
pre-commit run --all-filesUse clear and meaningful commit messages:
- Use the imperative mood ("Add feature" not "Added feature")
- Keep the first line under 72 characters
- Reference issues and PRs where appropriate
Example:
Add FP8 blockwise quantization kernel
- Implements per-token-group quantization for FP8
- Optimized for SM90 (Hopper) architecture
- Adds corresponding tests and benchmarks
Fixes #123
-
Update your fork with the latest upstream changes:
git fetch upstream git rebase upstream/main
-
Create a feature branch:
git checkout -b feature/your-feature-name
-
Make your changes following the coding standards
-
Add or update tests as necessary
-
Run the test suite locally:
make test -
Update documentation if needed (README, AGENTS.md, code comments)
-
Commit your changes with clear messages
-
Push to your fork:
git push origin feature/your-feature-name
-
Create a Pull Request on GitHub
- Maintainers will review your PR as soon as possible
- Address review comments by pushing additional commits
- Once approved, a maintainer will merge your PR
- Code follows the project's coding standards
- All pre-commit hooks pass
- Tests pass locally
- New tests added for new functionality
- Documentation updated
- Commit messages are clear and descriptive
If you have questions or need help, feel free to:
- Open a GitHub Discussion
- Join our community channels (if available)
Thank you for contributing to tf-kernel!