Skip to content

perf(api): skip GIL release for HMAC encode/decode (#124) - #130

Merged
ZhuchkaTriplesix merged 2 commits into
devfrom
issue/124-hmac-gil-release
Sep 28, 2026
Merged

ZhuchkaTriplesix merged 2 commits into
devfrom
issue/124-hmac-gil-release

Conversation

@ZhuchkaTriplesix

Copy link
Copy Markdown
Member

Problem

encode, encode_json, decode and decode_complete always call py.detach, releasing and reacquiring the GIL around every native call. For HS256/384/512 the sign/verify itself takes roughly a microsecond, so the release-and-reacquire cycle can cost as much as the operation for negligible concurrency benefit on that single call — and, as the benchmark below shows, under thread contention the cycle itself becomes a source of lock/scheduler overhead, not just per-call cost.

This is explicitly an "evaluate" issue: measure first, change only if it actually helps.

Fix

Added maybe_detach, which releases the GIL unless told to skip it, and skip it specifically for HMAC:

  • encode/encode_json: the resolved algorithm is already known.
  • decode/decode_complete: checked against the caller's algorithms allow-list before the token header is parsed — an HMAC-only allow-list already guarantees the verified algorithm is HMAC too, since a mixed-family allow-list is rejected elsewhere.

RSA, EC and EdDSA are untouched and keep releasing the GIL unconditionally.

Results (release build, this machine)

before after
single-threaded HS256 decode ~1.18 µs ~1.15 µs
8 threads continuously decoding the same HS256 token ~410k decode/sec ~790k decode/sec (~1.9× more)

The multi-threaded number was the interesting finding: I expected the change to be roughly throughput-neutral for multi-threaded workloads (per the issue's "does not regress" bar) and was surprised it nearly doubled — releasing and reacquiring a contended GIL on every ~1 µs call turns out to add real futex/scheduler overhead under load, on top of the per-call cost. Without an explicit release, CPython's own bytecode-tick-based switching yields far less often, which is exactly what keeps HMAC-only workloads fast here.

Confirmed the unaffected path still scales with threads (detach is still exercised for non-HMAC):

1 thread 8 threads
RS256 decode ~67k/sec ~339k/sec

Free-threaded Python

No correctness concern: the HMAC path never calls back into Python inside the (now possibly not released) section, so holding the GIL/attachment throughout is never a hazard on 3.13t/3.14t — it's strictly a choice not to release something that didn't need to be released for that closure's own safety.

Tests

Added test_hmac_encode_decode_are_thread_safe_without_gil_release (tests/test_encode_decode.py): 8 threads round-tripping their own payload through a shared secret, 200 iterations each, asserting no cross-talk.

Checklist

  • Benchmark with/without detach for HS256 encode/decode, single-threaded and multi-threaded (8 threads) — see above
  • Change applied because the single-threaded gain is measurable and multi-threaded throughput does not regress (it substantially improves)
  • cargo build/clippy --all-targets -D warnings/fmt --check/cargo test clean with both aws_lc_rs (default) and --no-default-features --features rust_crypto
  • Full pytest (271 passed, 1 skipped) and mypy clean
  • No behavioral change

Closes #124.

encode, encode_json, decode and decode_complete always called
py.detach, releasing and reacquiring the GIL around every native
call. For HS256/384/512 the sign/verify itself takes roughly a
microsecond, so under thread contention the release-and-reacquire
cycle can cost as much as the operation itself (or more, from
futex/scheduler overhead on the contended lock), for negligible
concurrency benefit on a call that short.

Add maybe_detach, which releases the GIL unless told to skip it, and
skip it specifically for HMAC: on encode, the resolved algorithm is
already known; on decode/decode_complete, checked against the
caller's algorithms allow-list before the token header is parsed
(an HMAC-only allow-list already guarantees the verified algorithm
is HMAC too, since a mixed-family allow-list is rejected). RSA, EC
and EdDSA keep releasing the GIL unconditionally.

Measured (release build):
- Single-threaded HS256 decode: ~1.18us -> ~1.15us.
- 8 threads continuously decoding the same HS256 token:
  ~410k -> ~790k decodes/sec (~1.9x more). The GIL release/reacquire
  cycle was itself a source of contention overhead under load, not
  just single-call cost.
- RS256 decode (unaffected path) still scales with threads:
  ~67k/sec on 1 thread, ~339k/sec on 8, confirming detach is still
  exercised for non-HMAC algorithms.

No behavioural change. Safe on free-threaded Python 3.13t/3.14t: the
HMAC path never calls back into Python, so holding the GIL/attachment
throughout is never a correctness hazard, only a choice not to
release it.

Closes #124
Regression test for the previous commit: with the GIL no longer
released around HS256 encode/decode, many threads round-tripping
their own payload through the same secret must still get back
exactly what they encoded, with no cross-talk between threads.
@ZhuchkaTriplesix
ZhuchkaTriplesix merged commit 81875f7 into dev Sep 28, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant