Skip to content

perf(lite): bound read allocations and fix shared-library exports - #1022

Merged
ajroetker merged 4 commits into
mainfrom
codex/lite-decoded-cache-budget
Oct 8, 2026
Merged

ajroetker merged 4 commits into
mainfrom
codex/lite-decoded-cache-budget

Conversation

@ajroetker

@ajroetker ajroetker commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Repeated local reads decoded and copied the same payloads, and duplicate result pins exhausted the owner budget even when they referred to one block. Impossible cache admissions could also evict useful blocks, and a compressed-prefix key miss caused a second block load. This follow-up to #1018 fixes those paths and bounds decoder working memory while preserving ordinary point-read ownership.

It also fixes #1023: libantfly exported internal symbols, including __dso_handle, causing aws-lc consumers to fail their direct constructor relocations. Shared-library exports now come from the public C header, with exact allowlists for Mach-O and ELF.

Changes:

  • Bound decoded-cache payload residency to 2 MiB by default; bypass blocks over 256 KiB. Deduplicate result pins before applying the 1 MiB / 64 distinct-block owner bounds. Reject impossible admissions before eviction, and keep metadata allocation failures optional.
  • Promote eligible prefix/prefix-snappy blocks after two misses, with bounded metadata and cooldown after denied admission. A confirmed prefix miss returns directly instead of reloading and promoting the block.
  • Share cold local-block decodes through 16 fixed flight records, including failure delivery and cleanup. Backend-locked callers decode independently to avoid waiting for a cache publisher that needs their mutex.
  • Reuse four decoder arenas, each retaining at most 64 KiB capacity, behind an 8 MiB default active working-byte gate and an enforced scratch-allocation cap. Exceptional oversized operations run alone. The pool also serves direct compressed point reads and owned fallback loaders.
  • Account scratch and uncached decoded payload leases in the new lsm.read_working_set resource slice, keeping existing resource IDs stable. Cache admission atomically transfers credit to the cache slice without a second host charge. Final result release returns evicted payload credit; a nonblocking reclaimer and shutdown fence preserve callback lifetime.
  • Borrow eligible local results in write batches, transfer existing owned allocations instead of copying again, and reuse batch metadata for up to 1,024 keys. Larger batches use temporary scratch. Mutable results remain owned before their captured view is released; result pins retain the transaction-wide limits across batches.
  • Honor transient/read-once cache policy for local scans and point promotion. Such reads can reuse an existing hot block without admitting their own cold blocks.
  • Separate scratch from output allocation in compressed decoders, validate declared lengths before allocation, and allocate decoded output exactly once.
  • Generate export allowlists from antfly.h and verify both leaked internals and missing public functions. macOS links the C API archive and its provider archives with an exported-symbol list; native builds use Apple ld and cross builds use ld64.lld (overridable with -Dmacos-linker). Preserve SDK, deployment version, frameworks, runtime paths, header padding, and ad-hoc signing. ELF uses an exact version script. Windows retains its original artifact/import-library handling.
  • Run export verification before installing shared libraries, and under C smoke validation. Constructor regression consumers link both an executable and another shared library; the arm64 fixture reproduces the exact arm64_adrp_lo12 failure against the unrestricted Zig dylib. Add policy/linkage tests to CI and LLVM inspection tools to release builders.

Measurements on aarch64 macOS, Debug, with deterministic fixtures:

  • 100 point reads after the first cold read: prefix backend allocations 500 → 5; prefix-snappy 604 → 5; encoded-block loads 100 → 1. A subsequent 100 warm reads allocate and load zero backend blocks.
  • 101 two-key read batches: retained pins 64 → 1 and owned values 74 → 0.
  • 100 warm two-key write batches, compared with the previous PR head 1ad9a16b18: backend allocations 206 → 0 and owned values 200 → 0, with one retained block pin and no additional writer scratch allocations after warmup.
  • Eight concurrent cold batch readers share one encoded-block load, verified for both success and injected failure.
  • 100 transient compressed point reads use 100 backend allocations for compact selected rows, for both codecs; scratch is reused and the hot cache is preserved.
  • A compressed-prefix Bloom false positive now uses one encoded-block load and admits no decoded block, versus two loads and one admission before.

Backend allocation counts exclude caller-owned point results and runtime metadata supplied by its separate allocator; writer scratch reuse is checked separately. Encoded loads may hit provider memory caches. These are allocation/lifetime fixtures, not a query-throughput or 2 GiB soak benchmark. Results and policy bounds are recorded in zig/bench/baselines/lite-decoded-cache-budget.json.

Validation:

  • Focused allocation, ownership, writer/probe, codec, resource, and concurrency suite: 230 passed, 2 skipped, zero failures or leaks.
  • Broader native storage suite: 853 passed, 21 skipped, zero failures or leaks.
  • WASM platform checks and embedded shared smoke test: passed.
  • Zig formatting, JSON parsing, and whitespace checks: passed.
  • Full macOS ReleaseSafe shared libraries linked by Apple ld and LLVM ld64.lld: each exposes exactly 98 header-declared functions, passes C smoke and executable/shared constructor regressions, and passes all 14 conformance cases.
  • Export-policy/linkage tests: 5 passed on macOS and 5 passed in a network-isolated Linux container.
  • Packaging/source-release/license tests: 18 passed.
  • Full Linux aarch64 GNU ReleaseSafe shared library: exactly 98 public functions, no internal exports.
  • Regression coverage includes impossible and denied admission, allocator failure cleanup, independent ordinary point-result ownership, eviction with pinned leases, mutable-result publication lifetime, shared decode failures, bounded decoder retention, oversized-work isolation, cache charge transfer, host reclamation under a held backend mutex, and shutdown racing callback registration.

origin/main was already integrated at 84dfbf5a95. The on-disk format and header-declared C ABI are unchanged. Undocumented internal functions are no longer exported.

Fixes #1023.

@ajroetker ajroetker changed the title perf(lite): bound decoded cache and reuse hot read payloads perf(lite): bound decoded cache and share local read scratch Oct 8, 2026
@ajroetker ajroetker changed the title perf(lite): bound decoded cache and share local read scratch perf(lite): bound read allocations and fix shared-library exports Oct 8, 2026
@ajroetker
ajroetker merged commit 8a8f338 into main Oct 8, 2026
1 of 2 checks passed
@ajroetker
ajroetker deleted the codex/lite-decoded-cache-budget branch October 8, 2026 20:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

libantfly.dylib exports ___dso_handle, libc/libm, compiler-rt, ObjC, and Metal internals; breaks linking with aws-lc

1 participant