Thank you for your interest in contributing to PromptDiff! ⚡
Whether you are reporting a bug, proposing a new feature, writing documentation, or submitting a pull request, we welcome your contributions.
PromptDiff requires Python 3.10+. We recommend using a virtual environment:
# 1. Fork and clone the repository
git clone https://github.com/latryee/promptdiff.git
cd promptdiff
# 2. Create and activate a virtual environment
python -m venv .venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
# 3. Install in editable mode with development dependencies
pip install --upgrade pip
pip install -e ".[dev,all]"New to PromptDiff or looking for an easy place to start? We actively curate bite-sized, accessible tasks labeled good first issue. Here are 5 great entry points:
- Add a New LLM Pricing Model Entry (
promptdiff/core/pricing.py):- Add input/output token pricing per 1M tokens for newly released models (e.g.
gemini-2.0-flash,claude-3-7-sonnet,deepseek-v3). - Add matching assertions in
tests/core/test_providers.py.
- Add input/output token pricing per 1M tokens for newly released models (e.g.
- Implement a Specialized Evaluator Metric (
promptdiff/evaluators/):- Create a lightweight evaluator checking output formatting rules (e.g., Markdown table validator, YAML schema verifier, or reading grade level scorer).
- Inherit from
BaseEvaluatorand register it inEvaluatorRegistry.
- Expand the Adversarial Red-Teaming Fuzzer (
promptdiff/security/fuzzer.py):- Add new jailbreak, indirect prompt injection, or system prompt exfiltration payload patterns to increase test coverage.
- CLI Output Polish & Terminal UX (
promptdiff/cli/):- Improve Rich table formatting, add colored status badges, or enhance error hints when user datasets contain missing fields.
- Add End-to-End Evaluation Recipes (
examples/recipes/):- Contribute a runnable evaluation recipe demonstrating prompt regression testing with frameworks like LangGraph, Instructor, or LiteLLM.
Before submitting a pull request, ensure that all tests, linters, and type checks pass:
pytest -vWe use Ruff for high-speed linting and formatting:
# Lint checks
ruff check .
# Auto-fix linting issues
ruff check --fix .
# Format checks
ruff format --check .
# Auto-format
ruff format .We maintain strict Mypy type safety:
mypy promptdiffWe dogfood PromptDiff on our own prompts — see .promptdiff-self-test/, prompt evolution from v1 to the optimizer-compressed version:
promptdiff test .promptdiff-self-test/prompts/system_v1.txt .promptdiff-self-test/prompts/system_v2.txt --inputs .promptdiff-self-test/testcases.jsonl --mock-
Create a branch for your work:
git checkout -b feat/my-new-feature # or git checkout -b fix/my-bugfix -
Commit your changes: We follow Conventional Commits:
feat(evaluator): add semantic perplexity scorerfix(cli): resolve windows terminal encoding bugdocs: improve pytest plugin quickstart exampletest: add unit test for council evaluator consensus
-
Verify locally: Ensure
pytest,ruff check, andmypyall pass cleanly. -
Submit your PR: Open a pull request against
main. Ensure your PR description explains the motivation, implementation details, and verification steps.
- Questions & Discussions: GitHub Discussions
- Bug Reports & Feature Requests: GitHub Issues