perf(api): skip GIL release for HMAC encode/decode (#124) - #130
Merged
Merged
Conversation
encode, encode_json, decode and decode_complete always called py.detach, releasing and reacquiring the GIL around every native call. For HS256/384/512 the sign/verify itself takes roughly a microsecond, so under thread contention the release-and-reacquire cycle can cost as much as the operation itself (or more, from futex/scheduler overhead on the contended lock), for negligible concurrency benefit on a call that short. Add maybe_detach, which releases the GIL unless told to skip it, and skip it specifically for HMAC: on encode, the resolved algorithm is already known; on decode/decode_complete, checked against the caller's algorithms allow-list before the token header is parsed (an HMAC-only allow-list already guarantees the verified algorithm is HMAC too, since a mixed-family allow-list is rejected). RSA, EC and EdDSA keep releasing the GIL unconditionally. Measured (release build): - Single-threaded HS256 decode: ~1.18us -> ~1.15us. - 8 threads continuously decoding the same HS256 token: ~410k -> ~790k decodes/sec (~1.9x more). The GIL release/reacquire cycle was itself a source of contention overhead under load, not just single-call cost. - RS256 decode (unaffected path) still scales with threads: ~67k/sec on 1 thread, ~339k/sec on 8, confirming detach is still exercised for non-HMAC algorithms. No behavioural change. Safe on free-threaded Python 3.13t/3.14t: the HMAC path never calls back into Python, so holding the GIL/attachment throughout is never a correctness hazard, only a choice not to release it. Closes #124
Regression test for the previous commit: with the GIL no longer released around HS256 encode/decode, many threads round-tripping their own payload through the same secret must still get back exactly what they encoded, with no cross-talk between threads.
3 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
encode,encode_json,decodeanddecode_completealways callpy.detach, releasing and reacquiring the GIL around every native call. For HS256/384/512 the sign/verify itself takes roughly a microsecond, so the release-and-reacquire cycle can cost as much as the operation for negligible concurrency benefit on that single call — and, as the benchmark below shows, under thread contention the cycle itself becomes a source of lock/scheduler overhead, not just per-call cost.This is explicitly an "evaluate" issue: measure first, change only if it actually helps.
Fix
Added
maybe_detach, which releases the GIL unless told to skip it, and skip it specifically for HMAC:encode/encode_json: the resolved algorithm is already known.decode/decode_complete: checked against the caller'salgorithmsallow-list before the token header is parsed — an HMAC-only allow-list already guarantees the verified algorithm is HMAC too, since a mixed-family allow-list is rejected elsewhere.RSA, EC and EdDSA are untouched and keep releasing the GIL unconditionally.
Results (release build, this machine)
decodeThe multi-threaded number was the interesting finding: I expected the change to be roughly throughput-neutral for multi-threaded workloads (per the issue's "does not regress" bar) and was surprised it nearly doubled — releasing and reacquiring a contended GIL on every ~1 µs call turns out to add real futex/scheduler overhead under load, on top of the per-call cost. Without an explicit release, CPython's own bytecode-tick-based switching yields far less often, which is exactly what keeps HMAC-only workloads fast here.
Confirmed the unaffected path still scales with threads (
detachis still exercised for non-HMAC):decodeFree-threaded Python
No correctness concern: the HMAC path never calls back into Python inside the (now possibly not released) section, so holding the GIL/attachment throughout is never a hazard on 3.13t/3.14t — it's strictly a choice not to release something that didn't need to be released for that closure's own safety.
Tests
Added
test_hmac_encode_decode_are_thread_safe_without_gil_release(tests/test_encode_decode.py): 8 threads round-tripping their own payload through a shared secret, 200 iterations each, asserting no cross-talk.Checklist
detachfor HS256 encode/decode, single-threaded and multi-threaded (8 threads) — see abovecargo build/clippy --all-targets -D warnings/fmt --check/cargo testclean with bothaws_lc_rs(default) and--no-default-features --features rust_cryptopytest(271 passed, 1 skipped) andmypycleanCloses #124.