Skip to content

Latest commit

 

History

History
277 lines (198 loc) · 8.1 KB

File metadata and controls

277 lines (198 loc) · 8.1 KB

Contributing to tf-kernel

Thank you for your interest in contributing to tf-kernel! This document provides guidelines and instructions for contributing to this project.

Table of Contents

Code of Conduct

This project adheres to the Contributor Covenant Code of Conduct. By participating, you are expected to uphold this code.

Getting Started

  1. Fork the repository on GitHub
  2. Clone your fork locally:
    git clone https://github.com/YOUR_USERNAME/TeleFuser.git
    cd TeleFuser/tf-kernel
  3. Set up the upstream remote:
    git remote add upstream https://github.com/Tele-AI/TeleFuser.git

Development Setup

Prerequisites

  • Python >= 3.10
  • CMake >= 3.26
  • CUDA Toolkit 12.8+
  • PyTorch == 2.11.0

Setup Development Environment

Option 1: Install development tools and build (Recommended)

Install development, testing, documentation, and linting tools without invoking a source installation through pip, then compile tf-kernel with Make:

# Install development tools
python -m pip install pytest pytest-cov sphinx sphinx-rtd-theme \
  sphinx-autodoc-typehints pre-commit isort ruff clang-format

# Install pre-commit hooks
pre-commit install

# Build and install the project (requires CUDA)
make build-auto PYTHON=/path/to/venv/bin/python

# Run tests
make test

Option 2: Install specific tool groups

If you only need specific dependencies:

# For running tests
python -m pip install pytest pytest-cov

# For building documentation
python -m pip install sphinx sphinx-rtd-theme sphinx-autodoc-typehints

# For code linting and formatting
python -m pip install pre-commit isort ruff clang-format

# Install pre-commit hooks (after installing lint dependencies)
pre-commit install

Do not use pip install ., pip install -e ., or pip install -e ".[dev]" for local tf-kernel source. Direct source installation fails during CMake configuration. Compilation and installation must go through a make build-* target, which installs the generated wheel.

Dependency Groups

Group Description Packages
dev All development dependencies Includes test, docs, and lint
test Testing dependencies pytest, pytest-cov
docs Documentation dependencies sphinx, sphinx-rtd-theme, sphinx-autodoc-typehints
lint Code formatting and linting pre-commit, black, isort, ruff

Note: clang-format for C++/CUDA formatting should be installed via system package manager:

# Ubuntu/Debian
sudo apt-get install clang-format

# macOS
brew install clang-format

Wheel Distribution Policy

Do not publish prebuilt tf-kernel wheels or a source distribution to PyPI or another public package index. Direct source installation is intentionally unsupported, so publishing an sdist would make installers attempt a build path that the project rejects. Do not add package-index installation instructions or a TeleFuser optional dependency.

Users may build a wheel on an explicitly provisioned CUDA/NVCC host and distribute it directly within a compatible, controlled environment. Before sharing an artifact:

make build-sm90 PYTHON=/path/to/venv/bin/python  # Select the required architecture.
make test-wheel PYTHON=/path/to/venv/bin/python
make test-smoke PYTHON=/path/to/venv/bin/python
sha256sum dist/*.whl

Record the source commit, package version, Python version, PyTorch version, PyTorch CUDA version, C++11 ABI, target SM family, CPU architecture, Linux/GLIBC baseline, wheel SHA-256, and test results with the artifact. Keep SM80, SM90, and SM100 wheels in separate artifact paths because architecture-specific builds currently produce the same wheel filename. Do not place multiple target-SM variants of one release in a single simple package index: pip cannot select a wheel from the target GPU architecture. Install a shared wheel by exact path or URL with --no-deps, then run python -m pip check and the installation smoke test on the target host.

How to Contribute

Reporting Bugs

Before creating a bug report, please check the existing issues to avoid duplicates.

When filing a bug report, please include:

  • System information: OS, CUDA version, GPU architecture
  • Python version and PyTorch version
  • Steps to reproduce the issue
  • Expected behavior vs actual behavior
  • Error messages or stack traces
  • Minimal code example that reproduces the issue

Suggesting Enhancements

Enhancement suggestions are tracked as GitHub issues. When creating an enhancement suggestion:

  • Use a clear and descriptive title
  • Provide a detailed description of the proposed feature
  • Explain why this enhancement would be useful
  • List some examples of how the feature would be used

Adding New Kernels

To add a new CUDA kernel:

  1. Implement the kernel in csrc/<category>/your_kernel.cu
  2. Declare the interface in include/tf_kernel_ops.h
  3. Register with PyTorch in csrc/common_extension.cc:
    m.def("your_kernel(Tensor input, Tensor! output) -> ()");
    m.impl("your_kernel", torch::kCUDA, &your_kernel);
  4. Update CMakeLists.txt: Add source file to the appropriate SOURCES list
  5. Create Python wrapper in tf_kernel/<category>.py
  6. Export in tf_kernel/__init__.py
  7. Add tests in tests/test_your_kernel.py
  8. Add benchmarks in benchmark/ (if applicable)

See AGENTS.md for more detailed technical information.

Coding Standards

C++/CUDA Code

  • Use clang-format with the provided .clang-format config (Google style based)
  • 2-space indentation
  • 120 column limit
  • Left pointer alignment (int* ptr not int *ptr)
# Format C++/CUDA files
make format

Python Code

  • isort: Import sorting
  • black: Code formatting (Disabled)
  • ruff: Linting (F821 rule, F401 disabled)
# Format Python files
make format

Pre-commit Hooks

All commits are checked by pre-commit hooks. You can run them manually:

pre-commit run --all-files

Commit Message Guidelines

Use clear and meaningful commit messages:

  • Use the imperative mood ("Add feature" not "Added feature")
  • Keep the first line under 72 characters
  • Reference issues and PRs where appropriate

Example:

Add FP8 blockwise quantization kernel

- Implements per-token-group quantization for FP8
- Optimized for SM90 (Hopper) architecture
- Adds corresponding tests and benchmarks

Fixes #123

Pull Request Process

  1. Update your fork with the latest upstream changes:

    git fetch upstream
    git rebase upstream/main
  2. Create a feature branch:

    git checkout -b feature/your-feature-name
  3. Make your changes following the coding standards

  4. Add or update tests as necessary

  5. Run the test suite locally:

    make test
  6. Update documentation if needed (README, AGENTS.md, code comments)

  7. Commit your changes with clear messages

  8. Push to your fork:

    git push origin feature/your-feature-name
  9. Create a Pull Request on GitHub

PR Review Process

  • Maintainers will review your PR as soon as possible
  • Address review comments by pushing additional commits
  • Once approved, a maintainer will merge your PR

PR Checklist

  • Code follows the project's coding standards
  • All pre-commit hooks pass
  • Tests pass locally
  • New tests added for new functionality
  • Documentation updated
  • Commit messages are clear and descriptive

Questions?

If you have questions or need help, feel free to:

Thank you for contributing to tf-kernel!