Skip to content

Release 0.8.0 - #133

Merged
ZhuchkaTriplesix merged 18 commits into
mainfrom
dev
Sep 28, 2026
Merged

ZhuchkaTriplesix merged 18 commits into
mainfrom
dev

Conversation

@ZhuchkaTriplesix

Copy link
Copy Markdown
Member

Summary

Release 0.8.0: version bump + changelog for the performance work already merged into dev (#120–#125).

[0.8.0] — 2026-09-28

Performance

  • RSA/RSA-PSS encode no longer re-parses the private key on every call.
    EncodingKey.from_rsa_pem now parses the DER-encoded key into an
    aws_lc_rs::signature::RsaKeyPair once, at construction time, and encode
    signs through that cached key directly. Previously jsonwebtoken::crypto::sign
    ran RsaKeyPair::from_der (including full RSA key validation) on every encode
    call, which dominated the cost of signing. Measured with a pre-built
    EncodingKey and a 2048-bit key: RS256 encode dropped from ~526 µs to
    ~201 µs per call (~2.6× faster), now within ~5% of cryptography's raw RSA
    sign with an equivalent pre-parsed key. decode, and encode/decode for
    HMAC, EC and EdDSA, are unaffected — those paths were already close to their
    theoretical floor. Only active with the default aws_lc_rs crypto backend;
    the rust_crypto feature (used for the Linux aarch64 wheel) is unchanged.
    A malformed RSA private key is now rejected by EncodingKey.from_rsa_pem
    itself instead of by the first encode call. (perf(keys): cache parsed aws-lc key objects instead of re-parsing per sign/verify #120)

  • EncodingKey / DecodingKey no longer cloned on every encode/decode
    call.
    Both pyclasses are now frozen, and the native encode, decode
    and decode_complete entry points borrow the underlying key material
    straight out of the Python object instead of cloning it (an owned copy of
    the DER/secret bytes, cloned again by jsonwebtoken's signer/verifier
    factory) on every call. HMAC secrets passed as raw str/bytes are
    unaffected — there is no persistent key object to borrow from in that case.
    No behavioural change. (perf(keys): borrow key material instead of cloning on every encode/decode #121)

  • Less Python-side work on the plain decode/decode_complete fast path.
    The _is_plain_decode argument check and the RFC 7797 detached_payload
    pre-check are now inlined into decode/decode_complete instead of going
    through a 9-argument and a keyword-argument function call. exp is only
    re-checked in Python when it is not a plain int: Rust already enforces
    exp > now with an integer clock on this path, which is the exact same
    predicate for an integer exp, so re-running it only cost a time.time()
    call with no behavioural difference; a float exp still gets the Python
    recheck, since Rust rounds a fractional value to the nearest second where
    PyJWT (and this check, for parity) truncates it, and the two can disagree
    right at the boundary. Measured (HS256, prebuilt token, int exp):
    wrapper overhead over the native call dropped from ~0.85 microseconds to
    ~0.57 microseconds (~33% less); with no time claims at all, ~0.37
    microseconds (~57% less). No behavioural change. (perf(api): trim Python overhead on the plain decode fast path #122)

  • Common claim/header names are interned instead of allocated per
    decode.
    json_to_bound (used by every decode and decode_complete
    call) used to allocate a fresh PyString for every dict key, including
    the same standard names — exp, iat, nbf, sub, aud, iss, jti,
    alg, typ, kid — on every single call. Those ten keys now come from
    pyo3::intern!, a per-key cache that both skips the repeated allocation
    and interns the string in CPython's own intern table, so a later
    payload.get("exp") on the Python side can hit the identity-comparison
    fast path for dict lookups. Any other key still gets an ordinary,
    uninterned PyString, unchanged from before. Measured on an 8-claim
    payload: native decode dropped from ~2.05 µs to ~1.92 µs (~6% less). No
    behavioural change. (perf(claims): intern common claim/header keys in json_to_bound #123)

  • HMAC encode/decode no longer release the GIL. encode,
    encode_json, decode and decode_complete used to call py.detach
    unconditionally, releasing and reacquiring the GIL around every native
    call. For HS256/384/512 the signing/verification itself takes roughly a
    microsecond, so under thread contention the release-and-reacquire cycle
    cost as much as the operation, or more: measured with 8 threads
    continuously decoding the same HS256 token, throughput went from ~410k to
    ~790k decodes/sec (~1.9× more) once the release was skipped, and
    single-threaded decode dropped from ~1.18 µs to ~1.15 µs. RSA, EC and
    EdDSA are unaffected — the check is on the resolved algorithm (encode)
    or on the caller's allow-list (decode, checked before the algorithm in
    the token header is known: an HMAC-only allow-list already guarantees the
    verified algorithm is HMAC too), and those algorithms still release the
    GIL, confirmed to keep scaling with threads (RS256 decode: ~67k/sec on 1
    thread, ~339k/sec on 8). No behavioural change; safe on free-threaded
    Python 3.13t/3.14t since the HMAC path never calls back into Python
    either way, so holding the GIL throughout is never a hazard, only a
    choice not to release it. (perf(api): evaluate skipping GIL release for HMAC operations #124)

  • Removed several small redundant allocations and re-parses on secondary
    decode/encode paths.
    None of these are on the main verified decode
    hot path (already addressed by earlier entries in this section); each is
    a modest, measured win on its own function:

    • get_unverified_header no longer copies the token into an owned
      String before py.detach (a borrowed &str is Ungil already; no
      'static bound requires the copy) and no longer runs its own
      split_compact_segments pre-check before parse_compact_header_json
      runs the exact same split internally. ~384 ns → ~352 ns.
    • decode_unverified used jsonwebtoken::dangerous::insecure_decode,
      which fully deserializes the header into jsonwebtoken's typed
      Header struct even though only .claims was ever read, and
      re-implements its own lenient segment split (silently misparsing a
      token with extra .s instead of rejecting it, unlike our own
      split_compact_segments) -- on top of the same redundant pre-check as
      get_unverified_header. Replaced with a single-pass helper that
      reuses our own strict split and only parses the header far enough to
      confirm it is a JSON object (matching get_unverified_header's own
      check) before discarding it. ~681 ns → ~595 ns.
    • decode_complete's unverified path (jws_parse_compact) computed and
      returned a header.payload "signing input" byte string that its only
      Python caller immediately discarded; it no longer computes it at all.
      ~787 ns → ~666 ns for the encode_json counterpart exercised by the
      same benchmark payload (the byte-building change below); the
      decode_complete(verify_signature=False) path itself is dominated by
      Python-side claim validation, so the native saving there is smaller
      (~2580 ns → ~2500 ns end to end).
    • decode_complete (verified path) decoded the signature segment's
      base64 a second time via a separate extract_signature_bytes call,
      which re-split the entire token from scratch to reach it. It now
      decodes the signature once, inline, right where the token is already
      split for verification -- removing the redundant full re-split;
      jsonwebtoken's crypto::verify still does its own internal base64
      decode of just the (small, bounded) signature segment, since it has
      no public entry point that accepts already-decoded signature bytes.
    • encode_json (and the RSA fast path shared with encode) built the
      header.payload.signature token through a chain of Engine::encode
      calls into throwaway Strings and two format!s, each copying
      everything built so far into a new allocation. It now encodes header
      and payload directly into one pre-sized String and appends the
      signature to the same buffer, so the whole token is built with (at
      most) one buffer growth instead of several full copies. ~787 ns →
      ~666 ns.

    No behavioural change, other than decode_unverified becoming slightly
    more lenient in one narrow, untested edge case: a token whose header is
    valid JSON but not a recognized alg name (e.g. {"alg": "made-up"})
    now decodes instead of raising, aligning it with get_unverified_header
    -- which already only required the header to be a JSON object -- rather
    than jsonwebtoken's stricter typed deserialization, which no other
    method in this library performs for an unverified decode. (perf(rust): remove redundant copies and re-parsing on secondary decode/encode paths #125)

ZhuchkaTriplesix and others added 18 commits September 28, 2026 13:08
jsonwebtoken::crypto::sign re-parses (and fully re-validates) the RSA
private key into an aws_lc_rs::signature::RsaKeyPair on every RS*/PS*
encode call, which was the dominant cost of RSA signing: RS256 encode
measured ~526us with a pre-built EncodingKey, ~2.7x the ~198us a raw
cryptography sign takes with an equivalent pre-parsed key.

EncodingKey.from_rsa_pem now parses the DER-encoded key into an
RsaKeyPair once, at construction time, and encode()/encode_json() sign
through that cached key directly via aws_lc_rs, bypassing
jsonwebtoken::crypto::sign for the RSA family. Every other algorithm
(HMAC, EC, EdDSA) and the decode/verify paths are unchanged; they were
already close to their measured floor.

Only active with the default aws_lc_rs crypto backend (behind
`#[cfg(feature = "aws_lc_rs")]`); the rust_crypto feature used for the
Linux aarch64 wheel keeps the previous jsonwebtoken-only path.

A malformed RSA private key is now rejected by EncodingKey.from_rsa_pem
itself instead of lazily by the first encode() call.

Measured (2048-bit key, pre-built EncodingKey): RS256 encode
526us -> 201us (2.6x), now within ~5% of raw cryptography sign (210us).
decode and all non-RSA algorithms are unaffected.

Closes #120
Add coverage for the cached-key path introduced in the previous
commit: repeated encode() calls on the same RSA EncodingKey (RS*/PS*)
must keep producing signatures that verify correctly, not just on the
first call, and a malformed RSA PEM must fail at
EncodingKey.from_rsa_pem construction rather than lazily at encode().
perf(keys): cache parsed RSA key pair instead of re-parsing per encode (#120)
EncodingKeyMaterial::encoding_key() and DecodingKeyMaterial::decoding_key()
returned an owned clone of the jsonwebtoken key on every call (an owned
copy of the DER bytes or HMAC secret), and jsonwebtoken's signer/verifier
factory cloned it again internally. For a prebuilt EncodingKey/DecodingKey
reused across many calls, that is pure overhead: the same key material
gets copied on every single encode/decode.

EncodingKey and DecodingKey are now `#[pyclass(frozen)]`. Frozen pyclasses
let pyo3 hand out `&T` via `Bound::get()` with no runtime borrow-flag
check and no clone, and that reference's lifetime is tied to the calling
scope rather than to a PyRef guard, so it can be threaded straight through
to `py.detach()`. encoding_key_from_py / decoding_key_from_py now return
a small BorrowedEncodingKey/BorrowedDecodingKey enum that either borrows
the key straight out of a frozen pyclass, or (for a raw HMAC secret with
no persistent key object) owns a freshly built one; both Deref to the
jsonwebtoken key type, so call sites are unchanged. The RSA fast path
added for #120 also drops its Arc<RsaKeyPair> clone in favor of a plain
borrow, since the frozen class covers the same lifetime need.

Raw str/bytes HMAC keys are unaffected: there is no persistent key object
to borrow from in that case, so an owned key is still built per call, same
as before.

No behavioural change; the classes had no &mut self pymethods to begin
with, so `frozen` has no Python-visible effect beyond enabling the borrow.

Measured (2048-bit RSA key, pre-built EncodingKey): RS256 encode
201us -> 187us (drops the now-unnecessary Arc clone from #120's cached
signing key), now ~1% over raw cryptography sign (190us).

Closes #121
perf(keys): borrow key material instead of cloning per encode/decode (#121)
For HS256 the Python wrapper cost about as much as the native decode
call itself (0.85us overhead on top of ~1.30us native). The fast path
in PyJWT.decode / decode_complete called _is_plain_decode (a
9-argument function) and _require_detached_payload_for_rfc7797 (a
keyword-argument call), and _validate_claims_default always called
time.time() and re-checked exp even though Rust had already validated
it with leeway=0.

- Inline the _is_plain_decode condition and the RFC 7797
  detached_payload pre-check directly into decode() and
  decode_complete(), removing the two function calls from the hot
  path. _require_detached_payload_for_rfc7797 stays as a function for
  the general (non-fast) path, which still needs its full signature.
- In _validate_claims_default, only re-check exp in Python when it is
  not a plain int. Rust already enforces exp > now with an integer
  clock on this path, which is the exact same predicate as the
  truncating check below for an integer exp. A float exp still needs
  the Python recheck: Rust rounds a fractional value to the nearest
  second while PyJWT (and this check, for parity) truncates it, so the
  two can disagree right at the boundary.
- time.time() is now only called when iat or a non-int exp is present.

Measured (HS256, prebuilt token): wrapper overhead over the native
call dropped from ~0.85us to ~0.57us with iat+int exp present (~33%
less), and to ~0.37us with no time claims at all (~57% less).

No behavioural change. Added a regression test pinning exp to a
fractional second (K + 0.6 for the current whole second K) to prove
the float recheck still runs: Rust rounds that up and would accept it,
but the Python truncating recheck must still reject it for PyJWT
parity.

Closes #122
Regression test for the exp re-check change in the previous commit:
a float exp pinned to K + 0.6 (K = current whole second) must still
raise ExpiredSignatureError via the Python recheck, even though Rust's
own rounding-based boundary check alone would accept it as not yet
expired.
…rhead

perf(api): trim Python overhead on the plain decode fast path (#122)
json_to_bound allocated a fresh PyString for every dict key on every
decode, including the same standard claim/header names every time:
exp, iat, nbf, sub, aud, iss, jti, alg, typ, kid.

Those ten keys now come from pyo3::intern!, which caches each key in
a call-site-local static rather than allocating it fresh. Interning
also registers the string in CPython's own intern table, so a
subsequent payload.get("exp") on the Python side (a str literal,
which CPython also interns) can hit the identity-comparison fast path
during dict lookup instead of a full string comparison. Any other key
still falls back to an ordinary, uninterned PyString, unchanged from
before.

Measured (native decode, 8-claim payload, release build):
~2.05us -> ~1.92us (~6% less).

No behavioural change.

Closes #123
Regression coverage for the previous commit: each standard claim name
and each standard header field must decode back as the exact same
interned str object pyo3 caches for it (identity, not just equality),
while a key outside that list must still decode correctly as an
ordinary, uninterned str.
perf(claims): intern common claim/header keys in json_to_bound (#123)
encode, encode_json, decode and decode_complete always called
py.detach, releasing and reacquiring the GIL around every native
call. For HS256/384/512 the sign/verify itself takes roughly a
microsecond, so under thread contention the release-and-reacquire
cycle can cost as much as the operation itself (or more, from
futex/scheduler overhead on the contended lock), for negligible
concurrency benefit on a call that short.

Add maybe_detach, which releases the GIL unless told to skip it, and
skip it specifically for HMAC: on encode, the resolved algorithm is
already known; on decode/decode_complete, checked against the
caller's algorithms allow-list before the token header is parsed
(an HMAC-only allow-list already guarantees the verified algorithm
is HMAC too, since a mixed-family allow-list is rejected). RSA, EC
and EdDSA keep releasing the GIL unconditionally.

Measured (release build):
- Single-threaded HS256 decode: ~1.18us -> ~1.15us.
- 8 threads continuously decoding the same HS256 token:
  ~410k -> ~790k decodes/sec (~1.9x more). The GIL release/reacquire
  cycle was itself a source of contention overhead under load, not
  just single-call cost.
- RS256 decode (unaffected path) still scales with threads:
  ~67k/sec on 1 thread, ~339k/sec on 8, confirming detach is still
  exercised for non-HMAC algorithms.

No behavioural change. Safe on free-threaded Python 3.13t/3.14t: the
HMAC path never calls back into Python, so holding the GIL/attachment
throughout is never a correctness hazard, only a choice not to
release it.

Closes #124
Regression test for the previous commit: with the GIL no longer
released around HS256 encode/decode, many threads round-tripping
their own payload through the same secret must still get back
exactly what they encoded, with no cross-talk between threads.
perf(api): skip GIL release for HMAC encode/decode (#124)
Several secondary decode/encode paths did work more than once for no
reason:

- get_unverified_header copied the token into an owned String before
  py.detach even though py.detach only requires the closure to be
  Ungil (Send), not 'static -- a borrowed &str already satisfies that.
  It also ran its own split_compact_segments pre-check right before
  parse_compact_header_json ran the exact same split internally.

- decode_unverified used jsonwebtoken::dangerous::insecure_decode,
  which fully deserializes the header into jsonwebtoken's typed
  Header struct even though only .claims is ever read, and
  re-implements its own lenient segment split that silently misparses
  a token with extra '.'s instead of rejecting it (unlike our own
  split_compact_segments) -- on top of the same redundant pre-check
  as get_unverified_header. Replaced with a new single-pass
  jws::parse_compact_claims_unverified that reuses the strict split,
  parses the header only far enough to confirm it's a JSON object
  (matching get_unverified_header's own check), then discards it, and
  parses only the payload.

- decode_verified_complete decoded the signature segment's base64 a
  second time via a separate extract_signature_bytes call, which
  re-split the *entire* token from scratch just to reach that one
  segment. verify_and_parse is now verify_and_parse_impl with an
  optional with_signature flag: when set, it decodes the signature
  once, inline, right where the token is already split for
  verification. jsonwebtoken's crypto::verify still does its own
  internal base64 decode of the (small, bounded) signature segment,
  since it has no public entry point that accepts pre-decoded bytes --
  removing the redundant full token re-split was the point, not the
  second signature-segment decode alone. extract_signature_bytes is
  now unused and removed.

- jws_parse_compact (decode_complete's unverified path) computed and
  returned a header.payload "signing input" byte string that its only
  Python caller (api_jwt.py) immediately discarded. It no longer
  computes it at all; parse_compact_jws's return type drops from a
  4-tuple to a 3-tuple (header, payload, signature).

- encode_json (and the RSA fast path it shares with encode via
  sign_compact_with_cached_rsa) built the final token through a chain
  of Engine::encode calls into throwaway Strings plus two format!s,
  each copying everything built so far into a new allocation. Both
  now encode header and payload directly into one pre-sized String
  (signing_input_string) and append the signature to the same buffer.

No behavioural change, except one narrow, previously-untested edge
case: decode_unverified now accepts a header that is valid JSON but
not a recognized alg name (e.g. {"alg": "made-up"}), since it no
longer deserializes into jsonwebtoken's typed Header struct. This
aligns it with get_unverified_header, which already only required the
header to be a JSON object; no other unverified-decode method in this
library enforces alg recognition, and unverified decode was never a
security boundary.

Measured (release build, HS256, small payload): get_unverified_header
~384ns -> ~352ns; decode_unverified ~681ns -> ~595ns; encode_json
~787ns -> ~666ns; decode_complete(verify_signature=False) end-to-end
~2580ns -> ~2500ns (mostly Python-side claim validation, so the native
saving is a smaller share of the total there).

Closes #125
Regression coverage for decode_unverified's move to a single-pass
claims-only parser: the header must still be rejected when it isn't a
JSON object at all, and must now be accepted (matching
get_unverified_header) when it's a well-formed JSON object with an
alg name jsonwebtoken's typed Header struct wouldn't recognize --
pinning the one intentional, narrow behaviour change from the
previous commit.
perf(rust): remove redundant copies and re-parsing on secondary paths (#125)
Bump version to 0.8.0 in pyproject.toml, rust/Cargo.toml and
python/oxyjwt/__init__.py, cut the Unreleased changelog section into
0.8.0, mirror it in docs-site/docs/changelog.md, and add
.github/RELEASE_NOTES_v0.8.0.md.
@ZhuchkaTriplesix
ZhuchkaTriplesix merged commit 88a77f2 into main Sep 28, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant