Problem
jemalloc 5.3.1 (the current submodule at 81034ce1) is missing a critical TSD/tcache initialization ordering fix. This causes production SIGSEGV when the tcache is marked as enabled before its bins backing storage is fully initialized.
Root Cause
In jemalloc's tsd_tcache_enabled_data_init() (src/tcache.c), the per-thread tcache is marked as enabled before tsd_tcache_data_init() completes. If the backing storage allocation fails, the return value is ignored, leaving the tcache in an enabled-but-uninitialized state. If a reentrant allocation occurs during the initialization window (e.g., from heap-profiling prof_tdata initialization or sampled backtrace collection), the fast path observes an enabled tcache with NULL cache-bin stack pointers and dereferences them.
Production Impact
We (GreptimeDB) experienced repeated SIGSEGV crashes across multiple days in production (greptime-ee 26.05.2.0, tikv-jemalloc-sys 0.6.0 with jemalloc 5.3.0-1-ge13ca993):
- SIGSEGV in
_rjem_je_arena_stats_merge (offsets 0x1645e239, 0x1645e300, 0x1645e364): triggered by Prometheus /metrics scrape → epoch.advance() → arena_stats_merge() traversing stale cache_bin_array_descriptor entries pointing to freed TLS/thread-stack memory.
- SIGSEGV in
calloc (offset 0x164547e1): dereferencing a NULL cache-bin stack pointer on a thread whose tcache was enabled but never properly initialized.
- #GP in libgcc unwinder (
libgcc_s.so.1+0x1583e): reading instruction data through a non-canonical unwind pointer during heap-profiling backtrace collection.
The crashes were probabilistic (spanning 4 days), consistent with a race condition requiring specific timing: heap profiling active (prof:true), heavy thread creation/destruction from I/O latency-induced blocking pool expansion, and memory pressure.
Evidence
We performed fault-injection testing on jemalloc 5.3.0 (same tcache code as the 5.3.0-1-ge13ca993 snapshot):
- Unpatched + injection: forcing tcache backing storage allocation failure every 5000th bootstrap → SIGSEGV within seconds, crash stack
free → arena_dalloc_large → cache_bin_full → tcache_bin_flush_stashed → san_check_stashed_ptrs, fault address NULL-derived (0xffffffffffffff60).
- Patched + injection (a056c20d semantics: disable tcache on init failure): 3.6M thread bootstraps, no crash.
- Unpatched +
tcache:false + injection: 2.5M threads, no crash.
- Unpatched, no injection: 5.4M threads, no crash (confirming the bug requires the allocation-failure trigger).
We also verified the exact fault offsets against the production binary (SHA256 685069fd..., Build ID cdcb1891...) — all addresses and instruction semantics match the described crash paths.
Upstream Status
The fix is already on jemalloc's dev branch:
- 54f22c83 — Initialize TSD tcache before enabling it (Farid Zakaria, 2026-07-06): moves
tsd_tcache_enabled_set() after tsd_tcache_data_init(), closing the reentrancy window.
- a056c20d — Handle tcache init failures gracefully (Carl Shapiro, 2026-03-02): disables tcache when initialization fails instead of leaving it in a broken state. This is included in 5.3.1.
However, 54f22c83 is NOT in 5.3.1 — it landed on dev 96 commits after 5.3.1 was tagged. The latest jemalloc release (5.3.1, 2026-04-13) does not include it.
Additionally, dev includes other TSD lifecycle fixes highly relevant to the stale descriptor crashes:
- fb5499aa9c — Handle jemalloc calls after TSD teardown: may address the stale
cache_bin_array_descriptor entries observed in arena_stats_merge.
- 61dc1da395 — Fix possible tcache corruption on fiber migration: another tcache corruption vector.
- 1e92317014 — Fix thread-exit TSD cleanup: TSD lifecycle improvement.
Proposal
Bump the jemalloc-sys/jemalloc submodule from 81034ce1 (5.3.1) to a dev branch commit that includes 54f22c83 and the related TSD fixes. The current dev HEAD is ff80bf2d (2026-09-10), which is 161 commits ahead of 5.3.1 and includes all of the above.
If bumping to dev HEAD is too aggressive, a smaller target like fb5499aa9c (which includes 54f22c83 + the TSD teardown fix) would also address the core issue.
We have a tested backport branch (GreptimeTeam/jemalloc#1) with 5.3.1 + 54f22c83 cherry-picked and verified, if a minimal-delta reference is useful.
References
Problem
jemalloc 5.3.1 (the current submodule at
81034ce1) is missing a critical TSD/tcache initialization ordering fix. This causes production SIGSEGV when the tcache is marked asenabledbefore its bins backing storage is fully initialized.Root Cause
In jemalloc's
tsd_tcache_enabled_data_init()(src/tcache.c), the per-thread tcache is marked as enabled beforetsd_tcache_data_init()completes. If the backing storage allocation fails, the return value is ignored, leaving the tcache in an enabled-but-uninitialized state. If a reentrant allocation occurs during the initialization window (e.g., from heap-profilingprof_tdatainitialization or sampled backtrace collection), the fast path observes an enabled tcache with NULL cache-bin stack pointers and dereferences them.Production Impact
We (GreptimeDB) experienced repeated SIGSEGV crashes across multiple days in production (greptime-ee 26.05.2.0, tikv-jemalloc-sys 0.6.0 with jemalloc 5.3.0-1-ge13ca993):
_rjem_je_arena_stats_merge(offsets0x1645e239,0x1645e300,0x1645e364): triggered by Prometheus/metricsscrape →epoch.advance()→arena_stats_merge()traversing stalecache_bin_array_descriptorentries pointing to freed TLS/thread-stack memory.calloc(offset0x164547e1): dereferencing a NULL cache-bin stack pointer on a thread whose tcache was enabled but never properly initialized.libgcc_s.so.1+0x1583e): reading instruction data through a non-canonical unwind pointer during heap-profiling backtrace collection.The crashes were probabilistic (spanning 4 days), consistent with a race condition requiring specific timing: heap profiling active (
prof:true), heavy thread creation/destruction from I/O latency-induced blocking pool expansion, and memory pressure.Evidence
We performed fault-injection testing on jemalloc 5.3.0 (same tcache code as the 5.3.0-1-ge13ca993 snapshot):
free → arena_dalloc_large → cache_bin_full → tcache_bin_flush_stashed → san_check_stashed_ptrs, fault address NULL-derived (0xffffffffffffff60).tcache:false+ injection: 2.5M threads, no crash.We also verified the exact fault offsets against the production binary (SHA256
685069fd..., Build IDcdcb1891...) — all addresses and instruction semantics match the described crash paths.Upstream Status
The fix is already on jemalloc's
devbranch:tsd_tcache_enabled_set()aftertsd_tcache_data_init(), closing the reentrancy window.However, 54f22c83 is NOT in 5.3.1 — it landed on
dev96 commits after 5.3.1 was tagged. The latest jemalloc release (5.3.1, 2026-04-13) does not include it.Additionally,
devincludes other TSD lifecycle fixes highly relevant to the stale descriptor crashes:cache_bin_array_descriptorentries observed inarena_stats_merge.Proposal
Bump the
jemalloc-sys/jemallocsubmodule from81034ce1(5.3.1) to adevbranch commit that includes 54f22c83 and the related TSD fixes. The currentdevHEAD isff80bf2d(2026-09-10), which is 161 commits ahead of 5.3.1 and includes all of the above.If bumping to
devHEAD is too aggressive, a smaller target likefb5499aa9c(which includes 54f22c83 + the TSD teardown fix) would also address the core issue.We have a tested backport branch (GreptimeTeam/jemalloc#1) with 5.3.1 + 54f22c83 cherry-picked and verified, if a minimal-delta reference is useful.
References