Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
18 commits
Select commit Hold shift + click to select a range
1ae274c
perf(keys): cache parsed RSA key pair instead of re-parsing per encode
ZhuchkaTriplesix Sep 28, 2026
234d2dd
test(algorithms): cover cached RSA signing key across repeated encodes
ZhuchkaTriplesix Sep 28, 2026
fae5bda
Merge pull request #126 from QueryaHub/issue/120-cached-parsed-keys
ZhuchkaTriplesix Sep 28, 2026
26f0d37
perf(keys): borrow key material instead of cloning per encode/decode
ZhuchkaTriplesix Sep 28, 2026
fae3e6c
Merge pull request #127 from QueryaHub/issue/121-borrow-keys
ZhuchkaTriplesix Sep 28, 2026
552b990
perf(api): trim Python overhead on the plain decode fast path
ZhuchkaTriplesix Sep 28, 2026
b7143b6
test(decode): cover float exp boundary parity on the fast path
ZhuchkaTriplesix Sep 28, 2026
0ec4c19
Merge pull request #128 from QueryaHub/issue/122-python-fast-path-ove…
ZhuchkaTriplesix Sep 28, 2026
024ee86
perf(claims): intern common claim/header keys in json_to_bound
ZhuchkaTriplesix Sep 28, 2026
a167bde
test(claims): cover claim/header key interning
ZhuchkaTriplesix Sep 28, 2026
f594932
Merge pull request #129 from QueryaHub/issue/123-intern-claim-keys
ZhuchkaTriplesix Sep 28, 2026
d908710
perf(api): skip GIL release for HMAC encode/decode
ZhuchkaTriplesix Sep 28, 2026
e23d6fb
test(encode): cover HMAC encode/decode concurrency
ZhuchkaTriplesix Sep 28, 2026
81875f7
Merge pull request #130 from QueryaHub/issue/124-hmac-gil-release
ZhuchkaTriplesix Sep 28, 2026
e5d4259
perf(rust): remove redundant copies and re-parsing on secondary paths
ZhuchkaTriplesix Sep 28, 2026
3aabf8a
test(decode): cover unverified decode header validation
ZhuchkaTriplesix Sep 28, 2026
771fa2e
Merge pull request #131 from QueryaHub/issue/125-minor-alloc-cleanup
ZhuchkaTriplesix Sep 28, 2026
6c08924
chore(release): prepare 0.8.0
ZhuchkaTriplesix Sep 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
35 changes: 35 additions & 0 deletions .github/RELEASE_NOTES_v0.8.0.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
# OxyJWT 0.8.0

**Beta** performance release — cached RSA signing keys, borrowed (no longer cloned) key material, a lighter Python decode fast path, interned claim/header names, and skipped GIL release on the HMAC hot path. No intentional breaking changes to the public `__all__` API.

## Highlights

### Performance

- **Cached RSA signing key** — `EncodingKey.from_rsa_pem` parses the DER key into an `aws_lc_rs::RsaKeyPair` once, at construction time; RS256 `encode` dropped from ~526 µs to ~201 µs (2.6× faster) with a pre-built key
- **No more per-call key cloning** — `EncodingKey` / `DecodingKey` are `frozen` pyclasses now; native `encode`/`decode` borrow the key material instead of cloning it
- **Lighter Python decode fast path** — inlined argument checks, and `exp` is only re-checked in Python when it isn't a plain `int`: wrapper overhead dropped from ~0.85 µs to ~0.57 µs (with `iat`) or ~0.37 µs (no time claims)
- **Interned claim/header names** (`exp`, `iat`, `nbf`, `sub`, `aud`, `iss`, `jti`, `alg`, `typ`, `kid`) — native `decode` on an 8-claim payload dropped from ~2.05 µs to ~1.92 µs
- **HMAC no longer releases the GIL** — HS256 decode throughput under 8-thread contention went from ~410k to ~790k decodes/sec; RSA/EC/EdDSA are unaffected and keep releasing the GIL
- Removed several redundant allocations/re-parses on secondary paths (`get_unverified_header`, `decode_unverified`, the unverified branch of `decode_complete`, `encode_json` token assembly)

### Behaviour change

- `decode_unverified` now accepts a header that is valid JSON but not a recognized `alg` name (e.g. `{"alg": "made-up"}`), matching `get_unverified_header`'s own check. Unverified decode was never a security boundary.

## Install

```bash
pip install oxyjwt==0.8.0
```

## Upgrade from 0.7.0

```bash
pip install -U oxyjwt
```

- No intentional breaking changes to public symbols.
- Successful encode/decode results are unchanged; only the `decode_unverified` edge case above differs.

See the full [changelog](https://github.com/QueryaHub/OxyJWT/blob/main/CHANGELOG.md#080--2026-09-28).
125 changes: 124 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,128 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

(No changes yet.)

## [0.8.0] — 2026-09-28

### Performance

- **RSA/RSA-PSS `encode` no longer re-parses the private key on every call.**
`EncodingKey.from_rsa_pem` now parses the DER-encoded key into an
`aws_lc_rs::signature::RsaKeyPair` once, at construction time, and `encode`
signs through that cached key directly. Previously `jsonwebtoken::crypto::sign`
ran `RsaKeyPair::from_der` (including full RSA key validation) on every `encode`
call, which dominated the cost of signing. Measured with a pre-built
`EncodingKey` and a 2048-bit key: RS256 `encode` dropped from ~526 µs to
~201 µs per call (~2.6× faster), now within ~5% of `cryptography`'s raw RSA
sign with an equivalent pre-parsed key. `decode`, and `encode`/`decode` for
HMAC, EC and EdDSA, are unaffected — those paths were already close to their
theoretical floor. Only active with the default `aws_lc_rs` crypto backend;
the `rust_crypto` feature (used for the Linux aarch64 wheel) is unchanged.
A malformed RSA private key is now rejected by `EncodingKey.from_rsa_pem`
itself instead of by the first `encode` call. (#120)
- **`EncodingKey` / `DecodingKey` no longer cloned on every `encode`/`decode`
call.** Both pyclasses are now `frozen`, and the native `encode`, `decode`
and `decode_complete` entry points borrow the underlying key material
straight out of the Python object instead of cloning it (an owned copy of
the DER/secret bytes, cloned again by `jsonwebtoken`'s signer/verifier
factory) on every call. HMAC secrets passed as raw `str`/`bytes` are
unaffected — there is no persistent key object to borrow from in that case.
No behavioural change. (#121)
- **Less Python-side work on the plain `decode`/`decode_complete` fast path.**
The `_is_plain_decode` argument check and the RFC 7797 `detached_payload`
pre-check are now inlined into `decode`/`decode_complete` instead of going
through a 9-argument and a keyword-argument function call. `exp` is only
re-checked in Python when it is not a plain `int`: Rust already enforces
`exp > now` with an integer clock on this path, which is the exact same
predicate for an integer `exp`, so re-running it only cost a `time.time()`
call with no behavioural difference; a `float` `exp` still gets the Python
recheck, since Rust rounds a fractional value to the nearest second where
PyJWT (and this check, for parity) truncates it, and the two can disagree
right at the boundary. Measured (HS256, prebuilt token, `int` `exp`):
wrapper overhead over the native call dropped from ~0.85 microseconds to
~0.57 microseconds (~33% less); with no time claims at all, ~0.37
microseconds (~57% less). No behavioural change. (#122)
- **Common claim/header names are interned instead of allocated per
decode.** `json_to_bound` (used by every `decode` and `decode_complete`
call) used to allocate a fresh `PyString` for every dict key, including
the same standard names — `exp`, `iat`, `nbf`, `sub`, `aud`, `iss`, `jti`,
`alg`, `typ`, `kid` — on every single call. Those ten keys now come from
`pyo3::intern!`, a per-key cache that both skips the repeated allocation
and interns the string in CPython's own intern table, so a later
`payload.get("exp")` on the Python side can hit the identity-comparison
fast path for dict lookups. Any other key still gets an ordinary,
uninterned `PyString`, unchanged from before. Measured on an 8-claim
payload: native `decode` dropped from ~2.05 µs to ~1.92 µs (~6% less). No
behavioural change. (#123)
- **HMAC `encode`/`decode` no longer release the GIL.** `encode`,
`encode_json`, `decode` and `decode_complete` used to call `py.detach`
unconditionally, releasing and reacquiring the GIL around every native
call. For HS256/384/512 the signing/verification itself takes roughly a
microsecond, so under thread contention the release-and-reacquire cycle
cost as much as the operation, or more: measured with 8 threads
continuously decoding the same HS256 token, throughput went from ~410k to
~790k decodes/sec (~1.9× more) once the release was skipped, and
single-threaded decode dropped from ~1.18 µs to ~1.15 µs. RSA, EC and
EdDSA are unaffected — the check is on the resolved algorithm (`encode`)
or on the caller's allow-list (`decode`, checked before the algorithm in
the token header is known: an HMAC-only allow-list already guarantees the
verified algorithm is HMAC too), and those algorithms still release the
GIL, confirmed to keep scaling with threads (RS256 decode: ~67k/sec on 1
thread, ~339k/sec on 8). No behavioural change; safe on free-threaded
Python 3.13t/3.14t since the HMAC path never calls back into Python
either way, so holding the GIL throughout is never a hazard, only a
choice not to release it. (#124)
- **Removed several small redundant allocations and re-parses on secondary
decode/encode paths.** None of these are on the main verified `decode`
hot path (already addressed by earlier entries in this section); each is
a modest, measured win on its own function:
- `get_unverified_header` no longer copies the token into an owned
`String` before `py.detach` (a borrowed `&str` is `Ungil` already; no
`'static` bound requires the copy) and no longer runs its own
`split_compact_segments` pre-check before `parse_compact_header_json`
runs the exact same split internally. ~384 ns → ~352 ns.
- `decode_unverified` used `jsonwebtoken::dangerous::insecure_decode`,
which fully deserializes the header into `jsonwebtoken`'s typed
`Header` struct even though only `.claims` was ever read, and
re-implements its own lenient segment split (silently misparsing a
token with extra `.`s instead of rejecting it, unlike our own
`split_compact_segments`) -- on top of the same redundant pre-check as
`get_unverified_header`. Replaced with a single-pass helper that
reuses our own strict split and only parses the header far enough to
confirm it is a JSON object (matching `get_unverified_header`'s own
check) before discarding it. ~681 ns → ~595 ns.
- `decode_complete`'s unverified path (`jws_parse_compact`) computed and
returned a `header.payload` "signing input" byte string that its only
Python caller immediately discarded; it no longer computes it at all.
~787 ns → ~666 ns for the `encode_json` counterpart exercised by the
same benchmark payload (the byte-building change below); the
`decode_complete(verify_signature=False)` path itself is dominated by
Python-side claim validation, so the native saving there is smaller
(~2580 ns → ~2500 ns end to end).
- `decode_complete` (verified path) decoded the signature segment's
base64 a second time via a separate `extract_signature_bytes` call,
which re-split the *entire* token from scratch to reach it. It now
decodes the signature once, inline, right where the token is already
split for verification -- removing the redundant full re-split;
`jsonwebtoken`'s `crypto::verify` still does its own internal base64
decode of just the (small, bounded) signature segment, since it has
no public entry point that accepts already-decoded signature bytes.
- `encode_json` (and the RSA fast path shared with `encode`) built the
`header.payload.signature` token through a chain of `Engine::encode`
calls into throwaway `String`s and two `format!`s, each copying
everything built so far into a new allocation. It now encodes header
and payload directly into one pre-sized `String` and appends the
signature to the same buffer, so the whole token is built with (at
most) one buffer growth instead of several full copies. ~787 ns →
~666 ns.

No behavioural change, other than `decode_unverified` becoming slightly
*more* lenient in one narrow, untested edge case: a token whose header is
valid JSON but not a recognized `alg` name (e.g. `{"alg": "made-up"}`)
now decodes instead of raising, aligning it with `get_unverified_header`
-- which already only required the header to be a JSON object -- rather
than `jsonwebtoken`'s stricter typed deserialization, which no other
method in this library performs for an *unverified* decode. (#125)

## [0.7.0] — 2026-08-26

Performance release. Verified `decode` is about **2.3× faster** and `encode` about
Expand Down Expand Up @@ -303,7 +425,8 @@ Initial alpha release.
- Mixed algorithm families are rejected for one decode call.
- In 0.1.0, `verify_signature=False` was rejected in `decode` (0.2.0 allows an explicit unverified path).

[Unreleased]: https://github.com/QueryaHub/OxyJWT/compare/v0.7.0...HEAD
[Unreleased]: https://github.com/QueryaHub/OxyJWT/compare/v0.8.0...HEAD
[0.8.0]: https://github.com/QueryaHub/OxyJWT/compare/v0.7.0...v0.8.0
[0.7.0]: https://github.com/QueryaHub/OxyJWT/compare/v0.6.0...v0.7.0
[0.6.0]: https://github.com/QueryaHub/OxyJWT/compare/v0.5.0...v0.6.0
[0.5.0]: https://github.com/QueryaHub/OxyJWT/compare/v0.4.0...v0.5.0
Expand Down
10 changes: 5 additions & 5 deletions RELEASING.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Releasing OxyJWT

This checklist is for maintainers publishing **0.7.x** (and later) to PyPI via the GitHub Actions [Release workflow](.github/workflows/release.yml).
This checklist is for maintainers publishing **0.8.x** (and later) to PyPI via the GitHub Actions [Release workflow](.github/workflows/release.yml).

## Before tagging

Expand Down Expand Up @@ -41,16 +41,16 @@ This checklist is for maintainers publishing **0.7.x** (and later) to PyPI via t
## Publish

1. Commit all release-prep changes on `main`.
2. Merge `dev` → `main`, then create and push an annotated tag (example for **0.7.0**):
2. Merge `dev` → `main`, then create and push an annotated tag (example for **0.8.0**):

```bash
git tag -a v0.7.0 -m "Release 0.7.0"
git push origin v0.7.0
git tag -a v0.8.0 -m "Release 0.8.0"
git push origin v0.8.0
```

3. The **Release** workflow runs full [CI](.github/workflows/ci.yml) via `workflow_call`, then builds wheels (Linux x86_64/aarch64, macOS, Windows) + sdist and publishes to PyPI only if CI passes (requires the `pypi` environment and [Trusted Publishing](https://docs.pypi.org/trusted-publishers/)).

4. On GitHub, create a **Release** from the tag. Use [`.github/RELEASE_NOTES_v0.7.0.md`](.github/RELEASE_NOTES_v0.7.0.md) or the **0.7.0** section in `CHANGELOG.md` as the release notes body.
4. On GitHub, create a **Release** from the tag. Use [`.github/RELEASE_NOTES_v0.8.0.md`](.github/RELEASE_NOTES_v0.8.0.md) or the **0.8.0** section in `CHANGELOG.md` as the release notes body.

## After release

Expand Down
72 changes: 72 additions & 0 deletions docs-site/docs/changelog.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,78 @@

(No changes yet.)

## 0.8.0 — 2026-09-28

Performance release: cached RSA signing keys, borrowed (no longer cloned) key
material, a lighter Python decode fast path, interned claim/header names, and
skipped GIL release on the HMAC hot path. No intentional breaking changes to
the public `__all__` surface. The API remains pre-1.0 (Beta). See
[Versioning](versioning.md) and [`SECURITY.md` on GitHub](https://github.com/QueryaHub/OxyJWT/blob/main/SECURITY.md).

### Upgrading from 0.7.0

```bash
pip install -U oxyjwt
```

- No intentional breaking changes to the public `__all__` surface.
- One edge-case behaviour change: `decode_unverified` now accepts a header
that is valid JSON but not a recognized `alg` name (for example
`{"alg": "made-up"}`), matching `get_unverified_header`'s own, looser
check. Unverified decode was never a security boundary, so this only
affects inspecting an untrusted token's claims without checking its
signature.

| Operation | Before | After | Change |
| --- | --- | --- | --- |
| RS256 `encode` (pre-built key) | 526 µs | 201 µs | 2.6× faster |
| `encode`/`decode` (pre-built key, clone removed) | — | — | no measurable regression, fewer allocations |
| `decode` wrapper overhead, `int` `exp` + `iat` | 0.85 µs | 0.57 µs | 33% less |
| `decode` wrapper overhead, no time claims | 0.85 µs | 0.37 µs | 57% less |
| native `decode`, 8-claim payload (key interning) | 2.05 µs | 1.92 µs | 6% less |
| HS256 `decode`, 8 threads (no GIL release) | ~410k/s | ~790k/s | 1.9× more |
| `get_unverified_header` | 384 ns | 352 ns | 8% less |
| `decode_unverified` | 681 ns | 595 ns | 13% less |
| `encode_json` | 787 ns | 666 ns | 15% less |

### Changed

- **RSA/RSA-PSS `encode` caches the parsed private key.** `EncodingKey.from_rsa_pem`
parses the DER-encoded key into an `aws_lc_rs` `RsaKeyPair` once, at
construction time, instead of `jsonwebtoken::crypto::sign` re-parsing (and
re-validating) it on every `encode` call. A malformed RSA private key is
now rejected by `EncodingKey.from_rsa_pem` itself instead of by the first
`encode` call. Only active with the default `aws_lc_rs` crypto backend;
the `rust_crypto` feature (Linux aarch64 wheel) is unchanged.
- **`EncodingKey` / `DecodingKey` are no longer cloned on every `encode` /
`decode` call.** Both pyclasses are now `frozen`, so the native entry
points borrow the underlying key material directly instead of cloning it.
Raw `str` / `bytes` HMAC secrets are unaffected.
- **Less Python-side work on the plain `decode` / `decode_complete` fast
path.** The argument checks that used to go through two extra function
calls are now inlined, and `exp` is only re-checked in Python when it is
not a plain `int` (Rust's own integer-clock check already covers that
case exactly; a `float` `exp` still gets the Python recheck for
boundary-rounding parity with PyJWT).
- **Common claim / header names are interned** (`exp`, `iat`, `nbf`, `sub`,
`aud`, `iss`, `jti`, `alg`, `typ`, `kid`) instead of allocated fresh on
every decode.
- **HMAC `encode` / `decode` no longer release the GIL.** HS256/384/512
sign/verify is fast enough that the release/reacquire cycle cost more
than the operation itself under thread contention. RSA, EC and EdDSA are
unaffected and keep releasing the GIL.
- Removed several redundant allocations and re-parses on secondary
decode/encode paths (`get_unverified_header`, `decode_unverified`, the
unverified branch of `decode_complete`, and `encode_json`'s token
assembly).

### Behaviour change

- `decode_unverified` now accepts a header that is valid JSON but is not a
recognized `alg` name, aligning it with `get_unverified_header`. No other
unverified-decode method enforces `alg` recognition, and unverified
decode was never a security boundary.

## 0.7.0 — 2026-08-26

Performance release. Verified `decode` is about **2.3× faster** and `encode` about
Expand Down
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ build-backend = "maturin"

[project]
name = "oxyjwt"
version = "0.7.0"
version = "0.8.0"
description = "High-performance Python JWT/JWS library backed by a Rust core (PyJWT drop-in replacement)."
readme = "README.md"
requires-python = ">=3.10"
Expand Down
2 changes: 1 addition & 1 deletion python/oxyjwt/__init__.py
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
"""OxyJWT public API (PyJWT-shaped module surface)."""

__version__ = "0.7.0"
__version__ = "0.8.0"

from ._oxyjwt import (
DecodingKey,
Expand Down
2 changes: 1 addition & 1 deletion python/oxyjwt/_oxyjwt.pyi
Original file line number Diff line number Diff line change
Expand Up @@ -92,7 +92,7 @@ def decode_verified_complete(

def jws_parse_compact(
token: str,
) -> tuple[bytes, dict[str, Any], bytes, bytes]: ...
) -> tuple[dict[str, Any], bytes, bytes]: ...


def get_unverified_header(token: str) -> dict[str, Any]: ...
Expand Down
Loading
Loading