Skip to content

fix(py): pin mimalloc v2 — v3 corrupts a co-resident pyarrow heap - #301

Merged
niko86 merged 4 commits into
mainfrom
fix/mimalloc-coresident-heap-corruption
Aug 13, 2026
Merged

niko86 merged 4 commits into
mainfrom
fix/mimalloc-coresident-heap-corruption

Conversation

@niko86

@niko86 niko86 commented Aug 13, 2026 •

Copy link
Copy Markdown
Owner

✅ Not stacked — branches off main (ec66197), full CI applies. Closes #294; re-points #122.

What it is

import pyarrow then import laterite was enough to corrupt pyarrow's buffers:

import pyarrow as pa      # pyarrow's bundled mimalloc initialises first
import laterite           # our cdylib brought a second mimalloc v3

vals = ["A" * 32, "B" * 32, "C" * 32, "D" * 32]
pa.table({"c": vals}).column(0).chunk(0).to_pylist()
# ['\x00' * 32, '\x00' * 32, 'CCCC…', 'DDDD…']   <- first 64 bytes zeroed

No laterite call. No Arrow FFI. 100% reproducible, and green with the pin.

pyarrow bundles its own mimalloc as its default memory pool, and libmimalloc-sys defaults to v3 from 0.1.47. Two co-resident mimalloc v3 instances hand out overlapping memory — microsoft/mimalloc#1327.

This is known upstream and already fixed on Arrow's side. apache/arrow GH-50428 — "pyarrow 24.0.0 regression: co-loading a native extension bundling mimalloc v3 SIGSEGVs in bundled mimalloc" — milestone 25.0.1, fixed by #50549. It independently names switching the extension's mimalloc copy to v2 as the workaround, which is exactly this PR.

So why pin at all, if pyarrow 25.0.1 fixes it? Because we do not control which pyarrow a user installs — [pyarrow] resolves to whatever their environment allows, and 24.0.0/25.0.0 are affected. GH-50428 also records that CPython 3.14 vendors its own mimalloc v3, so pyarrow is not the only co-residency partner in the process.

Measured across 50 array shapes on our v3 build: pyarrow 25.0.0 corrupts 44 of 50, pyarrow 25.0.1 corrupts 0 of 50. Under the v2 pin: 0 of 50 on both.

The diagnosis it replaces

#122 attributed this to an arrow-rs use-after-free consuming a pyarrow-produced ArrowArrayStream (apache/arrow-rs#10439) and landed a pl.from_pandas guard. That is not what is happening, which is why #294 kept reproducing through the guarded path.

claim measured
pandas-specific no — a raw pyarrow.Table in a live local corrupts identically
needs the Arrow FFI no — the minimal repro calls nothing
arrow-rs frees a foreign pointer no — an allocator shim watching pyarrow's buffer address never fired
only under a non-system allocator yes — and it needs both mimallocs

Cross-matrix (ours × pyarrow's pool): mimalloc × mimalloc → corrupt; mimalloc × system, system × mimalloc, system × system → clean. Corruption needs a string buffer >64 bytes with 2+ rows, in the first pyarrow allocations after our module loads.

Why not just drop the allocator

lib.rs sets mimalloc because the read hot-path is allocation-bound, and speed is the headline claim — so the cheap fix was the wrong one. On the 25 MB fixture (median of 7, ×3):

build read validate
v3 (was shipping) 115–118 ms 223–227 ms
v2 (this PR) 122–125 ms 229–233 ms
no mimalloc 150–152 ms 269–272 ms

v2 keeps ~18–19% read / ~14–15% validate over the system allocator. Removing mimalloc gives that back; the pin costs ~6% read against v3.

Scope

pyarrow only — polars and duckdb are clean, so the base pip install laterite was never exposed. Affected: [pyarrow], [compat,pyarrow], [all], and dev/CI (--all-extras).

laterite-node and laterite-cli take the same pin in a second commit. Neither is a proven exposure and neither is a fix: the addon loads no pyarrow but is dlopen'd beside addons we do not choose, and lat owns its whole process so the fault cannot reach it. Version hygiene — one workspace, one mimalloc major, so a future surface cannot inherit v3 by default.

Tests

The regression test runs in a subprocess: the fault depends on which allocator initialises first, so it is a property of process startup, and no in-process assertion can express it — by the time a test module imports, laterite is already loaded.

That is also the honest answer to "why did the #122 regression never catch this": neither of its assertions can. The texts match with or without the guard, and one native write cannot detect a corrupted heap. Its docstring now says so instead of claiming a guarantee it cannot check. I did not manufacture a second test at that seam — I tried, and both public pandas paths stay green even on the faulty build, so a test there would have been more of the same theatre.

  • test_pyarrow_buffers_survive_a_coresident_native_module — red on v3, green on v2, verified by swapping the two .so builds under the same test run.
  • Payload shape is load-bearing and commented as such, so nobody shrinks the fixture and silently disarms it.

Verification

The guard stays

build_ags4's pl.from_pandas guard is no longer what makes the pandas path memory-safe — the unguarded path is clean after the pin. But it is still load-bearing for a different reason: on a pyarrow-free [compat] install, pandas' __arrow_c_stream__ calls import_optional_dependency("pyarrow") and raises. Measured with pyarrow blocked: guarded → OK, unguarded → ImportError: Missing optional dependency 'pyarrow'. Comment corrected; retiring it is a separate, deliberate change.

Confirmed against the upstream reproducer

apache/arrow-rs#10439 has been withdrawn and closed, on evidence rather than inference. I rebuilt its reproducer verbatim and ran it at the versions originally reported, changing only the mimalloc major:

reproducer's Cargo.toml result
mimalloc = "0.1" (v3, as filed) SIGSEGV rc=139, 2/2
+ features = ["v2"] survives 500 iters, 3/3

arrow-rs identical across both. Worth recording two corrections to my own earlier reasoning:

  • It did not reproduce on my first attempt, because the resolver gave me pandas 3.0.5 instead of the reported 2.3.3. At the reported versions it segfaults reliably. A non-reproduction at the wrong versions would have been a bad reason to close it.
  • In that reproducer the arrow-rs consume call is necessary — a control keeping co-residency and heap churn but never calling consume() survives. I read that as arrow-rs being the workload that grows the Rust heap into the colliding region, not a defect in its consume path, and said so upstream so maintainers can disagree.

That prompted me to re-check the guard claim below under iteration rather than single-shot, since the upstream crash needs ~200 iterations: 500 iterations of the unguarded pandas → native emit survive on v2 and fail immediately on v3. The claim holds.

Not done here

  • Verified on macOS arm64 only. Whether Linux CI is exposed is untested.

`import pyarrow` followed by `import laterite` was enough to corrupt
pyarrow's buffers. No laterite call is required, no Arrow FFI is
involved, and the damage is silent until some later native call trips
over it — which is why the crash kept surfacing under an innocent test
name.

pyarrow bundles its own mimalloc as its default memory pool, and
libmimalloc-sys 0.1.49 defaults to v3. Two co-resident mimalloc v3
instances hand out overlapping memory (microsoft/mimalloc#1287, fixed
only on `dev3`, not in a stable release). apache/datafusion-python#1607
is the same fault and took the same pin.

Measured on the faulty build: corruption needs a string buffer over 64
bytes with 2+ rows, in the first pyarrow allocations after our module
loads — the first 64 bytes come back zeroed. Reproduced 100% of the
time; both orders and every consumer are clean after the pin.

Keeping the allocator matters, so the cheap fix was the wrong one.
Against the system allocator on the 25 MB fixture, v2 holds read ~18-19%
and validate ~14-15% faster; dropping mimalloc entirely gives that back
(read 122ms -> 150ms). v2 costs ~6% read against v3.

Scope: pyarrow only. polars and duckdb are unaffected, so the base
install was never exposed — only `[pyarrow]` / `[compat,pyarrow]` /
`[all]`, and dev/CI, which installs all extras.

The regression test runs in a subprocess: the fault depends on which
allocator initialises first, so it is a property of process startup that
no in-process assertion can express. That is why the existing #122
regression passed on every build that crashed; its docstring now says so
rather than claiming a guarantee it cannot check.

Also corrects the arrow-rs attribution in the `build_ags4` guard, which
sent readers to apache/arrow-rs#10439. The guard stays: it is no longer
what makes the pandas path memory-safe, but it is still what keeps that
path working on a pyarrow-free `[compat]` install, where pandas'
`__arrow_c_stream__` raises ImportError.
@codecov

codecov Bot commented Aug 13, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

niko86 added 3 commits August 13, 2026 02:05
Both carried a bare `mimalloc = "0.1"`, so both built the v3 that
corrupts a co-resident heap (microsoft/mimalloc#1287). Neither is a
proven exposure and neither is a fix:

- the Node addon loads no pyarrow, but it is dlopen'd into a process
  whose other native addons we do not choose — the same shape that bit
  the wheel;
- `lat` owns its whole process and loads no second allocator, so the
  fault cannot reach it at all. Pinned for version hygiene: one
  workspace, one mimalloc major, so a future surface cannot inherit v3
  by default.

`lat` builds and validates clean; laterite-cli's 47 tests pass.
#1287 is "mimalloc >= 3.3.0 causes segmentation faults when used from
multiple threads", a threading bug. It is not this fault, and it was
carried into every comment the previous two commits touched.

The co-residency report is microsoft/mimalloc#1327, and Arrow has the
same fault from its own side as apache/arrow GH-50428: a pyarrow 24.0.0
regression, milestone 25.0.1, whose accepted workaround is switching the
extension's mimalloc copy to v2 — the pin these commits make.

Two things that changes:

- pyarrow >=25.0.1 fixes it upstream, so the regression test only goes
  red on pyarrow 24.0.0-25.0.0. Measured: 44 of 50 shapes corrupt under
  v3 on 25.0.0, 0 of 50 on 25.0.1, 0 of 50 under v2 on both. Noted in
  the test, because a green run on a fixed pyarrow proves nothing.
- The pin stays regardless. We do not control which pyarrow a user
  installs, and CPython 3.14 vendors its own mimalloc v3, so pyarrow is
  not the only co-residency partner.
niko86 added a commit that referenced this pull request Aug 13, 2026
The page anchored its version-range claim on microsoft/mimalloc#1287. That is a
threading bug — "segmentation faults when used from multiple threads" — and our
fault is not a threading fault. The issue is real, and datafusion-python#1607
genuinely does root-cause to it, so it stays; what was wrong was using it as the
authority for a non-threading co-residency fault. It is now recorded as the
error it was, because the page exists to stop the next person citing a threading
bug for this.

The co-residency citation is microsoft/mimalloc#1327, which matches our
configuration closely: macOS arm64, a PyO3/abi3 cdylib linking mimalloc through
libmimalloc-sys, a discriminating version matrix, and a third instance we had
only inferred (CPython 3.14 vendors its own v3.3.2). apache/arrow GH-50428 is
the same fault from Arrow's side, fixed by apache/arrow#50549, and independently
names switching the extension to v2 — the same arm we reached from here.

Both of those crash at interpreter teardown, while ours is deterministic mid-run
zeroing, so they are cited as the same root-cause family and not as the same
symptom. Over-claiming that match is how the wrong issue got in.

Three status corrections: pyarrow 25.0.1 fixes it upstream, bounding exposure to
pyarrow 24.0.0-25.0.0; arrow-rs#10439 is closed, not open with no comments; and
the claim that the bug resists reproduction by iteration count was wrong. It
reproduces first try inside the fault window. The earlier non-reproduction was a
pyarrow version artifact, and pandas — briefly blamed — is not the variable.

Refs #122, #294, #301
@niko86
niko86 merged commit 050f1c6 into main Aug 13, 2026
18 checks passed
@niko86
niko86 deleted the fix/mimalloc-coresident-heap-corruption branch August 13, 2026 04:24
niko86 added a commit that referenced this pull request Aug 13, 2026
The branch was cut before PR #301 and sat while five more merged, so the page
described a tree that no longer exists in three places.

Two were live falsehoods, not staleness. The scope block said `laterite-node`
and `laterite-cli` "still take the v3 default" — true when written, made false
by the commit that pinned both alongside py; all three surfaces that install a
global allocator now pin v2, which is the whole point of the page. And the
pinned versions had moved underneath it: arrow 59.1.0 -> 59.2.0, PyO3
0.29.0 -> 0.29.2, polars 1.43.1 -> 1.43.2.

The third needed checking rather than editing. §2's chain is quoted from
arrow-array at tag 59.1.0, and the workspace now ships 59.2.0 — so either the
quotations still describe shipped code or they do not. `arrow-array/src/ffi.rs`
is byte-identical across the two tags (sha256 df4d434b…, 1968 lines both), so
they do; the page now says so and cites the comparison, rather than leaving a
reader to wonder why the scope and the verification disagree.

Deliberately NOT updated: "verified at tag 59.1.0" stays, because that is what
was read, and §7's measurement line keeps polars 1.43.1 because that is the venv
it was measured on. A provenance record that follows the tree is not a record.

`index.md` conflicted in the merge and was resolved by regenerating it, not by
picking hunks — it is generated, and `reindex.py --check` is the arbiter.
niko86 added a commit that referenced this pull request Aug 20, 2026
#448) (#476)

* ci(bench): measure what the macOS mimalloc pin costs Linux and Windows (#448)

The v2 pin (#294/#301) exists for a fault that is **macOS-specific**: mimalloc
v3 defaulted to Apple's fixed TLS slots 108/109, so two co-resident v3 instances
shared them regardless of symbol visibility. The pin is not platform-gated, so
Linux and Windows carry a macOS fix — and what it costs them has never been
measured. #448's own re-run list says to re-measure rather than quote the figure
from when the allocator landed.

`tools/bench/allocator_ab.py` builds `lat` three ways — v2, v3, and no
`#[global_allocator]` at all — and times them on a `forge scale` fixture, which
is synthesised from a seed rather than taken from a delivery, so it reproduces
anywhere and carries nobody's data.

No macOS leg, deliberately. The question is WITHIN-platform: same box, three
binaries. Comparing Linux-v3 against macOS-v2 would measure the machines.

Four things stop it reporting a number it has not earned.

- **The system arm is a positive control.** The perf ledger records
  mimalloc-vs-system as a real measured win, so if this harness cannot separate
  system from v2 it cannot resolve allocator effects at all — and "v2 and v3
  look the same" would mean the instrument is broken, not that the allocators
  are equivalent. That verdict is computed, not left to the reader.
- **A distinct-binary check.** If the feature swap fails to reach the compiler
  the three arms are one build measured three times; the run aborts on identical
  digests rather than reporting it.
- **An exact-match patcher.** A manifest that did not change still builds and
  still runs, reporting a number under the wrong label. A near-miss fails loudly.
- **A noise floor from two estimates, cleared by 2x.** A/B/A drift catches wander
  between arms; the widest within-arm spread catches a jittery box that lands
  both v2 passes together. An earlier revision called 1.5% "measurably faster"
  against a 1.0% floor — exactly the over-claim the harness exists to stop.

The spawn-per-rep overhead is a constant shared by every arm, so it shrinks the
measured gap rather than inflating it. The instrument under-claims, which is the
right bias for deciding whether to unpin.

Linux runs on the self-hosted pool because a dedicated box is the right
instrument for a timing comparison; Windows has no self-hosted runner and takes
`windows-2022`, which is noisier — and the harness's own floor is what says
whether those numbers are readable rather than a reader having to guess.

`workflow_dispatch`-only: it never fires on a push, a PR or a schedule, so it
cannot fail a build or enter the required set. Inventoried in the reliquary in
this same change, with removal keyed to the PR that records the measurement —
"temporary" otherwise survives by nobody remembering.

* ci(bench): point the Windows leg at the self-hosted VM, and stop assuming a shell

The Windows runner is registered on this repo now (`laterite-windows`), so the
leg becomes switchable rather than pinned to `windows-2022`: a `windows_runner`
dispatch input picks between them, defaulting to the github-hosted one so the
job still works if the VM is offline. The runner expression is one line on
purpose — a folded scalar keeps its newlines inside the `${{ }}`, and a bad
runner label does not error, it queues until the timeout.

Two corrections that `release.yml` had already paid for on this same VM, and
that this job would otherwise have rediscovered:

- **bash is not on PATH for the Windows runner service.** Four steps here used
  `shell: bash`. Split per-OS, as release.yml does for exactly this reason.
- **The default Windows shell is `pwsh`, which is not pre-installed.** A step
  with no `shell:` fails there. `powershell` is the built-in 5.1, and
  `-ExecutionPolicy Bypass` gets past the Win11 default of Restricted per-step
  rather than machine-wide.

Python comes from uv rather than from the box — it is not on PATH there, and
installing it would make the job depend on hand-configured machine state, which
the github-hosted fallback leg would not share. It is also the house pattern:
nothing here uses `setup-python`, and release.yml already proves the uv path on
this VM.

The rest is the harness taking back what the workflow could not say portably.
It now prints its own toolchain banner, writes its own `bench.log`, and appends
its own table to `GITHUB_STEP_SUMMARY` — so the `tee`/`sed` pipeline is gone
along with the third shell-shaped step, and a local run produces the same
artifacts as a CI one.

Verified by running it with `GITHUB_STEP_SUMMARY` set: log written, summary
written, and the verdict correct. At 4 reps on the 1 MB rung it reported *"the
instrument proved nothing"* — the control did not clear 2x the noise floor —
which is the guard working on a deliberately under-powered run rather than
reporting the 2.2% v3 gap sitting above it.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test(python): intermittent heap corruption in an in-process native call fails the coverage gate ~1 run in 4

1 participant