Release 0.8.0 - #133
Merged
Merged
Release 0.8.0#133
Conversation
jsonwebtoken::crypto::sign re-parses (and fully re-validates) the RSA private key into an aws_lc_rs::signature::RsaKeyPair on every RS*/PS* encode call, which was the dominant cost of RSA signing: RS256 encode measured ~526us with a pre-built EncodingKey, ~2.7x the ~198us a raw cryptography sign takes with an equivalent pre-parsed key. EncodingKey.from_rsa_pem now parses the DER-encoded key into an RsaKeyPair once, at construction time, and encode()/encode_json() sign through that cached key directly via aws_lc_rs, bypassing jsonwebtoken::crypto::sign for the RSA family. Every other algorithm (HMAC, EC, EdDSA) and the decode/verify paths are unchanged; they were already close to their measured floor. Only active with the default aws_lc_rs crypto backend (behind `#[cfg(feature = "aws_lc_rs")]`); the rust_crypto feature used for the Linux aarch64 wheel keeps the previous jsonwebtoken-only path. A malformed RSA private key is now rejected by EncodingKey.from_rsa_pem itself instead of lazily by the first encode() call. Measured (2048-bit key, pre-built EncodingKey): RS256 encode 526us -> 201us (2.6x), now within ~5% of raw cryptography sign (210us). decode and all non-RSA algorithms are unaffected. Closes #120
Add coverage for the cached-key path introduced in the previous commit: repeated encode() calls on the same RSA EncodingKey (RS*/PS*) must keep producing signatures that verify correctly, not just on the first call, and a malformed RSA PEM must fail at EncodingKey.from_rsa_pem construction rather than lazily at encode().
perf(keys): cache parsed RSA key pair instead of re-parsing per encode (#120)
EncodingKeyMaterial::encoding_key() and DecodingKeyMaterial::decoding_key() returned an owned clone of the jsonwebtoken key on every call (an owned copy of the DER bytes or HMAC secret), and jsonwebtoken's signer/verifier factory cloned it again internally. For a prebuilt EncodingKey/DecodingKey reused across many calls, that is pure overhead: the same key material gets copied on every single encode/decode. EncodingKey and DecodingKey are now `#[pyclass(frozen)]`. Frozen pyclasses let pyo3 hand out `&T` via `Bound::get()` with no runtime borrow-flag check and no clone, and that reference's lifetime is tied to the calling scope rather than to a PyRef guard, so it can be threaded straight through to `py.detach()`. encoding_key_from_py / decoding_key_from_py now return a small BorrowedEncodingKey/BorrowedDecodingKey enum that either borrows the key straight out of a frozen pyclass, or (for a raw HMAC secret with no persistent key object) owns a freshly built one; both Deref to the jsonwebtoken key type, so call sites are unchanged. The RSA fast path added for #120 also drops its Arc<RsaKeyPair> clone in favor of a plain borrow, since the frozen class covers the same lifetime need. Raw str/bytes HMAC keys are unaffected: there is no persistent key object to borrow from in that case, so an owned key is still built per call, same as before. No behavioural change; the classes had no &mut self pymethods to begin with, so `frozen` has no Python-visible effect beyond enabling the borrow. Measured (2048-bit RSA key, pre-built EncodingKey): RS256 encode 201us -> 187us (drops the now-unnecessary Arc clone from #120's cached signing key), now ~1% over raw cryptography sign (190us). Closes #121
perf(keys): borrow key material instead of cloning per encode/decode (#121)
For HS256 the Python wrapper cost about as much as the native decode call itself (0.85us overhead on top of ~1.30us native). The fast path in PyJWT.decode / decode_complete called _is_plain_decode (a 9-argument function) and _require_detached_payload_for_rfc7797 (a keyword-argument call), and _validate_claims_default always called time.time() and re-checked exp even though Rust had already validated it with leeway=0. - Inline the _is_plain_decode condition and the RFC 7797 detached_payload pre-check directly into decode() and decode_complete(), removing the two function calls from the hot path. _require_detached_payload_for_rfc7797 stays as a function for the general (non-fast) path, which still needs its full signature. - In _validate_claims_default, only re-check exp in Python when it is not a plain int. Rust already enforces exp > now with an integer clock on this path, which is the exact same predicate as the truncating check below for an integer exp. A float exp still needs the Python recheck: Rust rounds a fractional value to the nearest second while PyJWT (and this check, for parity) truncates it, so the two can disagree right at the boundary. - time.time() is now only called when iat or a non-int exp is present. Measured (HS256, prebuilt token): wrapper overhead over the native call dropped from ~0.85us to ~0.57us with iat+int exp present (~33% less), and to ~0.37us with no time claims at all (~57% less). No behavioural change. Added a regression test pinning exp to a fractional second (K + 0.6 for the current whole second K) to prove the float recheck still runs: Rust rounds that up and would accept it, but the Python truncating recheck must still reject it for PyJWT parity. Closes #122
Regression test for the exp re-check change in the previous commit: a float exp pinned to K + 0.6 (K = current whole second) must still raise ExpiredSignatureError via the Python recheck, even though Rust's own rounding-based boundary check alone would accept it as not yet expired.
…rhead perf(api): trim Python overhead on the plain decode fast path (#122)
json_to_bound allocated a fresh PyString for every dict key on every
decode, including the same standard claim/header names every time:
exp, iat, nbf, sub, aud, iss, jti, alg, typ, kid.
Those ten keys now come from pyo3::intern!, which caches each key in
a call-site-local static rather than allocating it fresh. Interning
also registers the string in CPython's own intern table, so a
subsequent payload.get("exp") on the Python side (a str literal,
which CPython also interns) can hit the identity-comparison fast path
during dict lookup instead of a full string comparison. Any other key
still falls back to an ordinary, uninterned PyString, unchanged from
before.
Measured (native decode, 8-claim payload, release build):
~2.05us -> ~1.92us (~6% less).
No behavioural change.
Closes #123
Regression coverage for the previous commit: each standard claim name and each standard header field must decode back as the exact same interned str object pyo3 caches for it (identity, not just equality), while a key outside that list must still decode correctly as an ordinary, uninterned str.
perf(claims): intern common claim/header keys in json_to_bound (#123)
encode, encode_json, decode and decode_complete always called py.detach, releasing and reacquiring the GIL around every native call. For HS256/384/512 the sign/verify itself takes roughly a microsecond, so under thread contention the release-and-reacquire cycle can cost as much as the operation itself (or more, from futex/scheduler overhead on the contended lock), for negligible concurrency benefit on a call that short. Add maybe_detach, which releases the GIL unless told to skip it, and skip it specifically for HMAC: on encode, the resolved algorithm is already known; on decode/decode_complete, checked against the caller's algorithms allow-list before the token header is parsed (an HMAC-only allow-list already guarantees the verified algorithm is HMAC too, since a mixed-family allow-list is rejected). RSA, EC and EdDSA keep releasing the GIL unconditionally. Measured (release build): - Single-threaded HS256 decode: ~1.18us -> ~1.15us. - 8 threads continuously decoding the same HS256 token: ~410k -> ~790k decodes/sec (~1.9x more). The GIL release/reacquire cycle was itself a source of contention overhead under load, not just single-call cost. - RS256 decode (unaffected path) still scales with threads: ~67k/sec on 1 thread, ~339k/sec on 8, confirming detach is still exercised for non-HMAC algorithms. No behavioural change. Safe on free-threaded Python 3.13t/3.14t: the HMAC path never calls back into Python, so holding the GIL/attachment throughout is never a correctness hazard, only a choice not to release it. Closes #124
Regression test for the previous commit: with the GIL no longer released around HS256 encode/decode, many threads round-tripping their own payload through the same secret must still get back exactly what they encoded, with no cross-talk between threads.
perf(api): skip GIL release for HMAC encode/decode (#124)
Several secondary decode/encode paths did work more than once for no
reason:
- get_unverified_header copied the token into an owned String before
py.detach even though py.detach only requires the closure to be
Ungil (Send), not 'static -- a borrowed &str already satisfies that.
It also ran its own split_compact_segments pre-check right before
parse_compact_header_json ran the exact same split internally.
- decode_unverified used jsonwebtoken::dangerous::insecure_decode,
which fully deserializes the header into jsonwebtoken's typed
Header struct even though only .claims is ever read, and
re-implements its own lenient segment split that silently misparses
a token with extra '.'s instead of rejecting it (unlike our own
split_compact_segments) -- on top of the same redundant pre-check
as get_unverified_header. Replaced with a new single-pass
jws::parse_compact_claims_unverified that reuses the strict split,
parses the header only far enough to confirm it's a JSON object
(matching get_unverified_header's own check), then discards it, and
parses only the payload.
- decode_verified_complete decoded the signature segment's base64 a
second time via a separate extract_signature_bytes call, which
re-split the *entire* token from scratch just to reach that one
segment. verify_and_parse is now verify_and_parse_impl with an
optional with_signature flag: when set, it decodes the signature
once, inline, right where the token is already split for
verification. jsonwebtoken's crypto::verify still does its own
internal base64 decode of the (small, bounded) signature segment,
since it has no public entry point that accepts pre-decoded bytes --
removing the redundant full token re-split was the point, not the
second signature-segment decode alone. extract_signature_bytes is
now unused and removed.
- jws_parse_compact (decode_complete's unverified path) computed and
returned a header.payload "signing input" byte string that its only
Python caller (api_jwt.py) immediately discarded. It no longer
computes it at all; parse_compact_jws's return type drops from a
4-tuple to a 3-tuple (header, payload, signature).
- encode_json (and the RSA fast path it shares with encode via
sign_compact_with_cached_rsa) built the final token through a chain
of Engine::encode calls into throwaway Strings plus two format!s,
each copying everything built so far into a new allocation. Both
now encode header and payload directly into one pre-sized String
(signing_input_string) and append the signature to the same buffer.
No behavioural change, except one narrow, previously-untested edge
case: decode_unverified now accepts a header that is valid JSON but
not a recognized alg name (e.g. {"alg": "made-up"}), since it no
longer deserializes into jsonwebtoken's typed Header struct. This
aligns it with get_unverified_header, which already only required the
header to be a JSON object; no other unverified-decode method in this
library enforces alg recognition, and unverified decode was never a
security boundary.
Measured (release build, HS256, small payload): get_unverified_header
~384ns -> ~352ns; decode_unverified ~681ns -> ~595ns; encode_json
~787ns -> ~666ns; decode_complete(verify_signature=False) end-to-end
~2580ns -> ~2500ns (mostly Python-side claim validation, so the native
saving is a smaller share of the total there).
Closes #125
Regression coverage for decode_unverified's move to a single-pass claims-only parser: the header must still be rejected when it isn't a JSON object at all, and must now be accepted (matching get_unverified_header) when it's a well-formed JSON object with an alg name jsonwebtoken's typed Header struct wouldn't recognize -- pinning the one intentional, narrow behaviour change from the previous commit.
perf(rust): remove redundant copies and re-parsing on secondary paths (#125)
Bump version to 0.8.0 in pyproject.toml, rust/Cargo.toml and python/oxyjwt/__init__.py, cut the Unreleased changelog section into 0.8.0, mirror it in docs-site/docs/changelog.md, and add .github/RELEASE_NOTES_v0.8.0.md.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Release 0.8.0: version bump + changelog for the performance work already merged into
dev(#120–#125).[0.8.0] — 2026-09-28
Performance
RSA/RSA-PSS
encodeno longer re-parses the private key on every call.EncodingKey.from_rsa_pemnow parses the DER-encoded key into anaws_lc_rs::signature::RsaKeyPaironce, at construction time, andencodesigns through that cached key directly. Previously
jsonwebtoken::crypto::signran
RsaKeyPair::from_der(including full RSA key validation) on everyencodecall, which dominated the cost of signing. Measured with a pre-built
EncodingKeyand a 2048-bit key: RS256encodedropped from ~526 µs to~201 µs per call (~2.6× faster), now within ~5% of
cryptography's raw RSAsign with an equivalent pre-parsed key.
decode, andencode/decodeforHMAC, EC and EdDSA, are unaffected — those paths were already close to their
theoretical floor. Only active with the default
aws_lc_rscrypto backend;the
rust_cryptofeature (used for the Linux aarch64 wheel) is unchanged.A malformed RSA private key is now rejected by
EncodingKey.from_rsa_pemitself instead of by the first
encodecall. (perf(keys): cache parsed aws-lc key objects instead of re-parsing per sign/verify #120)EncodingKey/DecodingKeyno longer cloned on everyencode/decodecall. Both pyclasses are now
frozen, and the nativeencode,decodeand
decode_completeentry points borrow the underlying key materialstraight out of the Python object instead of cloning it (an owned copy of
the DER/secret bytes, cloned again by
jsonwebtoken's signer/verifierfactory) on every call. HMAC secrets passed as raw
str/bytesareunaffected — there is no persistent key object to borrow from in that case.
No behavioural change. (perf(keys): borrow key material instead of cloning on every encode/decode #121)
Less Python-side work on the plain
decode/decode_completefast path.The
_is_plain_decodeargument check and the RFC 7797detached_payloadpre-check are now inlined into
decode/decode_completeinstead of goingthrough a 9-argument and a keyword-argument function call.
expis onlyre-checked in Python when it is not a plain
int: Rust already enforcesexp > nowwith an integer clock on this path, which is the exact samepredicate for an integer
exp, so re-running it only cost atime.time()call with no behavioural difference; a
floatexpstill gets the Pythonrecheck, since Rust rounds a fractional value to the nearest second where
PyJWT (and this check, for parity) truncates it, and the two can disagree
right at the boundary. Measured (HS256, prebuilt token,
intexp):wrapper overhead over the native call dropped from ~0.85 microseconds to
~0.57 microseconds (~33% less); with no time claims at all, ~0.37
microseconds (~57% less). No behavioural change. (perf(api): trim Python overhead on the plain decode fast path #122)
Common claim/header names are interned instead of allocated per
decode.
json_to_bound(used by everydecodeanddecode_completecall) used to allocate a fresh
PyStringfor every dict key, includingthe same standard names —
exp,iat,nbf,sub,aud,iss,jti,alg,typ,kid— on every single call. Those ten keys now come frompyo3::intern!, a per-key cache that both skips the repeated allocationand interns the string in CPython's own intern table, so a later
payload.get("exp")on the Python side can hit the identity-comparisonfast path for dict lookups. Any other key still gets an ordinary,
uninterned
PyString, unchanged from before. Measured on an 8-claimpayload: native
decodedropped from ~2.05 µs to ~1.92 µs (~6% less). Nobehavioural change. (perf(claims): intern common claim/header keys in json_to_bound #123)
HMAC
encode/decodeno longer release the GIL.encode,encode_json,decodeanddecode_completeused to callpy.detachunconditionally, releasing and reacquiring the GIL around every native
call. For HS256/384/512 the signing/verification itself takes roughly a
microsecond, so under thread contention the release-and-reacquire cycle
cost as much as the operation, or more: measured with 8 threads
continuously decoding the same HS256 token, throughput went from ~410k to
~790k decodes/sec (~1.9× more) once the release was skipped, and
single-threaded decode dropped from ~1.18 µs to ~1.15 µs. RSA, EC and
EdDSA are unaffected — the check is on the resolved algorithm (
encode)or on the caller's allow-list (
decode, checked before the algorithm inthe token header is known: an HMAC-only allow-list already guarantees the
verified algorithm is HMAC too), and those algorithms still release the
GIL, confirmed to keep scaling with threads (RS256 decode: ~67k/sec on 1
thread, ~339k/sec on 8). No behavioural change; safe on free-threaded
Python 3.13t/3.14t since the HMAC path never calls back into Python
either way, so holding the GIL throughout is never a hazard, only a
choice not to release it. (perf(api): evaluate skipping GIL release for HMAC operations #124)
Removed several small redundant allocations and re-parses on secondary
decode/encode paths. None of these are on the main verified
decodehot path (already addressed by earlier entries in this section); each is
a modest, measured win on its own function:
get_unverified_headerno longer copies the token into an ownedStringbeforepy.detach(a borrowed&strisUngilalready; no'staticbound requires the copy) and no longer runs its ownsplit_compact_segmentspre-check beforeparse_compact_header_jsonruns the exact same split internally. ~384 ns → ~352 ns.
decode_unverifiedusedjsonwebtoken::dangerous::insecure_decode,which fully deserializes the header into
jsonwebtoken's typedHeaderstruct even though only.claimswas ever read, andre-implements its own lenient segment split (silently misparsing a
token with extra
.s instead of rejecting it, unlike our ownsplit_compact_segments) -- on top of the same redundant pre-check asget_unverified_header. Replaced with a single-pass helper thatreuses our own strict split and only parses the header far enough to
confirm it is a JSON object (matching
get_unverified_header's owncheck) before discarding it. ~681 ns → ~595 ns.
decode_complete's unverified path (jws_parse_compact) computed andreturned a
header.payload"signing input" byte string that its onlyPython caller immediately discarded; it no longer computes it at all.
~787 ns → ~666 ns for the
encode_jsoncounterpart exercised by thesame benchmark payload (the byte-building change below); the
decode_complete(verify_signature=False)path itself is dominated byPython-side claim validation, so the native saving there is smaller
(~2580 ns → ~2500 ns end to end).
decode_complete(verified path) decoded the signature segment'sbase64 a second time via a separate
extract_signature_bytescall,which re-split the entire token from scratch to reach it. It now
decodes the signature once, inline, right where the token is already
split for verification -- removing the redundant full re-split;
jsonwebtoken'scrypto::verifystill does its own internal base64decode of just the (small, bounded) signature segment, since it has
no public entry point that accepts already-decoded signature bytes.
encode_json(and the RSA fast path shared withencode) built theheader.payload.signaturetoken through a chain ofEngine::encodecalls into throwaway
Strings and twoformat!s, each copyingeverything built so far into a new allocation. It now encodes header
and payload directly into one pre-sized
Stringand appends thesignature to the same buffer, so the whole token is built with (at
most) one buffer growth instead of several full copies. ~787 ns →
~666 ns.
No behavioural change, other than
decode_unverifiedbecoming slightlymore lenient in one narrow, untested edge case: a token whose header is
valid JSON but not a recognized
algname (e.g.{"alg": "made-up"})now decodes instead of raising, aligning it with
get_unverified_header-- which already only required the header to be a JSON object -- rather
than
jsonwebtoken's stricter typed deserialization, which no othermethod in this library performs for an unverified decode. (perf(rust): remove redundant copies and re-parsing on secondary decode/encode paths #125)