POC: Memory profiler allocation labels - #62649
Conversation
|
Review requested:
|
| // This happens when --experimental-async-context-frame is not set on | ||
| // Node.js 22, causing all contexts to map to Smi::zero() (address 0). | ||
| if (cped.IsEmpty() || cped->IsUndefined()) return; | ||
| uintptr_t addr = node::GetLocalAddress(cped); |
There was a problem hiding this comment.
Storing in binding data by the CPED address won't work at all. Because all AsyncLocalStorage contexts are combined into a single AsyncContextFrame map, any changes to any contexts will change what this value is, even if the particular store you are interested in has not changed at all within that map frame.
You would need to have V8 capture the CPED value at the time of the sample and store that on the heap profile itself alongside the samples, then use that actual AsyncContextFrame instance to look up what the corresponding data was in that frame for the label store.
8743634 to
302bebe
Compare
szegedi
left a comment
There was a problem hiding this comment.
Very interesting, I had to take a look since I was mentioned in the PR description itself 😄. I generally like the direction this is going in, long term we can probably replace Datadog's heap profiler that directly wraps V8 heap profiler with this and reduce our maintenance surface. This solution indeed looks like it can only be implemented in Node.js itself, and not as an add-on since the v8::ArrayBuffer::Allocator instance is a global Isolate::CreateParams setting so it's controlled by the embedder, that is, Node.js.
| v8::Local<v8::Value> context = | ||
| v8_isolate->GetContinuationPreservedEmbedderData(); | ||
| if (!context.IsEmpty() && !context->IsUndefined()) { | ||
| sample->cped.Reset(v8_isolate, context); |
There was a problem hiding this comment.
Can't you get the value associated with the AsyncLocalStorage from the AsyncContextFrame here and only store that? Essentially the additional step from BuildSamples. You'll be retaining in memory all ALSes this ACF references as keys, some might have large retained set themselves. If you can safely call GetContinuationPreservedEmbedderData (which creates a v8::Local) I'd think you can also safely get a local to the ALS key from its global, and call v8::Map::Get on context too? (Or direct V8 hashmap reading like I suggested in that other comment in ProfilingArrayBufferAllocator::FindCurrentLabels)
There was a problem hiding this comment.
Great point about memory retention. You're right that storing the entire CPED keeps all ALS stores alive as long as the sample exists, not just the labels store 💣 💥
OrderedHashMap::FindEntry is a great suggestion!
| // BackingStore::Allocate inside the ArrayBuffer constructor). | ||
| // Use AsArray() which reads the internal backing store directly without | ||
| // calling JS builtins, then iterate entries by identity comparison. | ||
| v8::Local<v8::Array> entries = frame->AsArray(); |
There was a problem hiding this comment.
What's the memory requirement of this AsArray call? It sounds like it'd have to construct a whole new array?
You're lucky that you have a whole embedded copy of V8, so you can use existing internals, something like this roughly sketched might work:
#include "src/objects/js-collection.h"
#include "src/objects/ordered-hash-table.h"
// Given a v8::Local<v8::Map>, get to the internal table:
i::Tagged<i::JSMap> js_map = *Utils::OpenDirectHandle(*frame);
i::Tagged<i::OrderedHashMap> table = i::Cast<i::OrderedHashMap>(js_map->table());
// no-JS lookup in the table:
i::InternalIndex entry = table->FindEntry(isolate, *Utils::OpenDirectHandle(*als_key));
if (entry.is_found()) {
i::Tagged<i::Object> value = table->ValueAt(entry);
// go back from Tagged to Local:
v8::Local<v8::Value> val = Utils::ToLocal(i::direct_handle(value, i_isolate));
// use val as before...
}There was a problem hiding this comment.
yeah AsArray() allocates a full JS Array and copies the Map backing store. I moved away from AsArray() to Map::Get() in an unpushed revision, but your OrderedHashMap::FindEntry approach would be even better.
Since we're already modifying V8 source in deps/v8/src/profiler/ and have access to internals from src/ as well, I can adopt this pattern in both locations:
SampleObject(allocation time) — extract ALS value, store as Global on SampleProfilingArrayBufferAllocator::TrackAllocate— same pattern for ArrayBuffer tracking
The read-time callback in GetAllocationProfile then receives the already-extracted flat array, making it trivial (just string conversion).
| @@ -11,6 +11,11 @@ | |||
| #include <unordered_set> | |||
| #include <vector> | |||
|
|
|||
| #ifdef V8_HEAP_PROFILER_SAMPLE_LABELS | |||
There was a problem hiding this comment.
So I presume this'll need upstreaming, right?
There was a problem hiding this comment.
yeah, I'm hoping if we can show that Nodejs would definitely use this feature they'd be more open to accepting it
rudolf
left a comment
There was a problem hiding this comment.
Very interesting, I had to take a look since I was mentioned in the PR description itself 😄. I generally like the direction this is going in, long term we can probably replace Datadog's heap profiler that directly wraps V8 heap profiler with this and reduce our maintenance surface.
@szegedi Thanks for popping by! Your thread sparked this idea and made me think maybe it's not all that hard (at least for memory profiling, sounds like CPU profiles might be a different beast).
| v8::Local<v8::Value> context = | ||
| v8_isolate->GetContinuationPreservedEmbedderData(); | ||
| if (!context.IsEmpty() && !context->IsUndefined()) { | ||
| sample->cped.Reset(v8_isolate, context); |
There was a problem hiding this comment.
Great point about memory retention. You're right that storing the entire CPED keeps all ALS stores alive as long as the sample exists, not just the labels store 💣 💥
OrderedHashMap::FindEntry is a great suggestion!
| // BackingStore::Allocate inside the ArrayBuffer constructor). | ||
| // Use AsArray() which reads the internal backing store directly without | ||
| // calling JS builtins, then iterate entries by identity comparison. | ||
| v8::Local<v8::Array> entries = frame->AsArray(); |
There was a problem hiding this comment.
yeah AsArray() allocates a full JS Array and copies the Map backing store. I moved away from AsArray() to Map::Get() in an unpushed revision, but your OrderedHashMap::FindEntry approach would be even better.
Since we're already modifying V8 source in deps/v8/src/profiler/ and have access to internals from src/ as well, I can adopt this pattern in both locations:
SampleObject(allocation time) — extract ALS value, store as Global on SampleProfilingArrayBufferAllocator::TrackAllocate— same pattern for ArrayBuffer tracking
The read-time callback in GetAllocationProfile then receives the already-extracted flat array, making it trivial (just string conversion).
| @@ -11,6 +11,11 @@ | |||
| #include <unordered_set> | |||
| #include <vector> | |||
|
|
|||
| #ifdef V8_HEAP_PROFILER_SAMPLE_LABELS | |||
There was a problem hiding this comment.
yeah, I'm hoping if we can show that Nodejs would definitely use this feature they'd be more open to accepting it
302bebe to
da6af28
Compare
| // label array) instead of the full CPED avoids retaining all ALS | ||
| // stores for the lifetime of the sample. Labels are resolved from | ||
| // this at read time (in GetAllocationProfile) via the callback. | ||
| Global<Value> label_value; |
There was a problem hiding this comment.
This means we are storing the label value per sample, even when many samples share the same label set. I think this is likely expensive in memory and in GetAllocationProfile()
IMO it's better to store a small uint32_t label_id = 0 on each sample instead, and keep a shared map for labels. I believe that with this GetAllocationProfile() could resolve labels once per unique context rather than once per sample.
There was a problem hiding this comment.
I agree in principle but wonder if the memory overhead is really an issue. We only retain samples and memory that have survived GC. Starting with per sample overhead: ~120 B base + ~24 B label handle slot = ~144 B. And assuming Poisson sampling with 512kb intervals:
- 200 MB live heap would have ~400 retained samples with 58 KB / 0.029% sample overhead (48 KB base, 10 KB labels)
- 1GB live heap would have ~2000 samples with 288kb sample overhead (240KB base, 48KB labels)
There's additional overhead like the shared ALS array but that's a function of number of unique lables not a per sample overhead so would be present in both designs. Feels worth profiling/benchmarking to make sure my math adds up, but the label_id solution comes at the cost of additional complexity for a small memory gain.
It makes a lot of sense to add a local cache during GetAllocationProfile to avoid the expensive callbacks for labels we have already resolved. I think that can make a noticeable difference to CPU overhead.
There was a problem hiding this comment.
We only retain samples and memory that have survived GC
Probably not if you call it like:
hp->StartSamplingHeapProfiler(
64, 16,
static_cast<v8::HeapProfiler::SamplingFlags>(
v8::HeapProfiler::kSamplingIncludeObjectsCollectedByMinorGC |
v8::HeapProfiler::kSamplingIncludeObjectsCollectedByMajorGC
)
);
There was a problem hiding this comment.
@IlyasShabi I'm stuck trying to measure this. The labels work adds bookkeeping that lives in C++ heap so process.memoryUsage().rss and v8.getHeapStatistics().used_heap_size don't directly capture it.
What I've tried so far:
- v8.getHeapStatistics().used_global_handles_size surprisingly doesn't track Sample::label_value Globals
- RSS deltas across 3-iteration matrix runs at 30 s each has a noise band of +/-60 MB
Is there a methodology you'd use to measure this kind of work?
There was a problem hiding this comment.
Did some measurement work on this. With a realistic HTTP workload (~2,500 rps, ~750 MB/s V8 alloc churn), 30 s with includeCollectedObjects: true at64 KB sampling interval, ~300K retained samples: aggregate libc-malloc shows the labels feature adds 6-9 MB on top of ~29 MB of profiler-only bookkeeping. Below 5% of the underlying sampler cost at this workload, scales linearly with sample count.
However, while running the load test I hit a segfault in ProfilingArrayBufferAllocator::TrackFree: V8's ArrayBufferSweeper calls it on a background worker thread, and the per-allocation Global label_value runs ~Global off-thread which trips the node->IsInUse() CHECK in GlobalHandles::NodeSpace::Release.
The fix it I think we need the same machinery as what you suggested for Sample::label_value so will take a stab at that
| if (sampleInterval !== undefined) validateUint32(sampleInterval, 'sampleInterval', true); | ||
| if (stackDepth !== undefined) validateUint32(stackDepth, 'stackDepth'); | ||
| if (options !== undefined) validateObject(options, 'options'); | ||
| return _startSamplingHeapProfiler(sampleInterval, stackDepth, options); |
There was a problem hiding this comment.
We should call ensureHeapProfileLabelsALS() before this line to make sure the ALS store exists and its key is registered before sampling begins.
Oh yeah, I could probably talk for hours about CPU sample labeling :-) Reading labels to associate with samples is the tricky part with CPU profiles. Since it happens in a signal handler, we we can't use basically any V8 APIs, as even getting a There's also some other fun details, like needing to have a ringbuffer for samples around as you can't allocate memory in the signal handler. And we're also using Associating labels in that ringbuffer with samples is also somewhat tricky, we do it by correlating monotonic clock values – we invoke As I said, I can probably go on for at least an hour. I might need to write a talk :-) |
GetAllocationProfile() invoked the labels callback once per sample, even when many samples shared the same ALS value (typical for any withHeapProfileLabels scope holding multiple allocations). Add a per-call cache keyed by ALS value Address. Empty resolutions are also cached to avoid re-invoking the callback for values that legitimately resolve to nothing. Cache misses on GC-moved objects are correctness safe; the callback runs again. Benchmark numbers in benchmark/v8/heap-profiler-labels-resolution.js. Refs: nodejs#62649 (comment) Signed-off-by: Rudolf Meijering <skaapgif@gmail.com>
Yeah, this is why I wanted to have V8 itself capture CPED pointer state in each sample internally so that could be read back as actual local values later when you process the profile tree. Never got around to that before I changed teams though. |
|
As far as design decisions go, have you considered not committing to having a set of string labels, but have a more generic facility for associating any JS value with samples? You wouldn't have to deal with string serialization and would just need to pass around a shareable global reference to a value ( This way you only implement a very generic facility in Node.js/V8 for associating values with samples, without restricting what those values can be. E.g. an OTEL tracing implementation might want to use a bigint for trace ID. An implementation might want to use numeric values for some labels. And so on. |
|
This pull request has been marked as stale due to 90 days of inactivity. |
Attach an opaque label to each heap profiler allocation sample so a profile can attribute memory to application context. Add LabelInternTable, which refcounts Global<Value> keyed on identity hash and hands out uint32_t ids, and a label_id field on AllocationProfile::Sample. The sampler captures the value held in ContinuationPreservedEmbedderData at allocation time and interns it; callers resolve the id back to the value through the profiler. The table is owned by HeapProfiler so it outlives the sampler, which holds a reference. Intern and Lookup require the isolate main thread; Release is safe from any thread, because the embedder frees tracked allocations from a background thread, and queues work that must run on the main thread. All of it sits behind V8_HEAP_PROFILER_SAMPLE_LABELS so builds that do not define the macro see no change, including no change to the layout of AllocationProfile::Sample. Signed-off-by: Rudolf Meijering <skaapgif@gmail.com>
Wire the V8 label machinery into Node. The v8 binding stores an AsyncLocalStorage instance as the label key, arms it for the duration of a labels session, and resolves label ids back into frozen label objects when a profile is read. Add ProfilingArrayBufferAllocator, which wraps the ArrayBuffer allocator to attribute off-heap memory to the same labels and reports it as externalBytes. It is installed only while a labels session is running; otherwise the allocate and free paths cost one relaxed atomic load. Frees arrive on V8's ArrayBufferSweeper background thread, so the enabled flag is the sentinel and the isolate pointer is written once per session and never cleared. Sessions carry a generation counter so a handle cannot read or stop a session it does not own, which is reachable because the V8 sampler is a single shared resource that the inspector can stop out of band. Define V8_HEAP_PROFILER_SAMPLE_LABELS wherever CPED is enabled, including in common.gypi so that native addons compiling against the shipped headers see the same struct layout as libnode. Signed-off-by: Rudolf Meijering <skaapgif@gmail.com>
Add a labels option to v8.startHeapProfile, and getAllocationProfile() on the returned handle so a labelled profile can be read while profiling continues. The profile reports samples with their labels and per-label external memory; stop() keeps returning the DevTools JSON string unchanged. Labels are set with v8.withHeapProfileLabels(labels, fn) for a scope, or v8.setHeapProfileLabels(labels) for the current context, and propagate through async work because they live in an AsyncLocalStorage. Values are validated as strings when set, so resolving a label later cannot run user code. worker.startHeapProfile rejects the labels option rather than accepting and ignoring it; labels work inside a worker that profiles itself. Signed-off-by: Rudolf Meijering <skaapgif@gmail.com>
Cover the label machinery in C++ and JavaScript. The cctests exercise the intern table directly, including refcounting, cross-thread release, the deferred free queue, and the destruction order that lets the table outlive the sampler. The JS tests cover label propagation across await boundaries and into workers, attribution accuracy against a known allocation ratio, external memory accounting, sessions that the inspector stops out of band, a poisoned Object.prototype, and the option gating that decides whether samples carry labels at all. Signed-off-by: Rudolf Meijering <skaapgif@gmail.com>
Document the labels option on v8.startHeapProfile, the getAllocationProfile() method on the handle, and the two functions that set labels. Document the limitations: labelled data is only available through getAllocationProfile() and never from stop(), external memory accounting needs Node's own ArrayBuffer allocator, small pooled Buffers are attributed to whichever label triggered the pool refill, and addons built outside node-gyp must define V8_HEAP_PROFILER_SAMPLE_LABELS to see the same struct layout as libnode. Signed-off-by: Rudolf Meijering <skaapgif@gmail.com>
Measure the cost of running with labels: sampling with and without labels enabled, label resolution when a profile is read, and two HTTP benchmarks that label by route. The two-server HTTP benchmark builds a DB response for the row count it was asked for, including one passed on the command line. Building only the two default sizes makes the server fall back to the 1000-row response, so a larger row count silently measures the smaller workload. Signed-off-by: Rudolf Meijering <skaapgif@gmail.com>
a6391ec to
7be5070
Compare
This is a POC for initial feedback. If we can get alignment within Node.js I could try to contribute the V8 changes upstream.
Summary
Adds the ability to tag sampling heap profiler allocations with string labels that propagate through async context (via CPED). This enables attributing memory usage to specific HTTP routes, tenants, or operations.
V8 changes
The V8 surface is intentionally minimal and avoids committing to string types.
AllocationProfile::Sample::label_id(uint32_t, behindV8_HEAP_PROFILER_SAMPLE_LABELS) — a compact 4-byte identifier instead of avector<pair<string,string>>.Sampleremains a plain aggregate.LabelInternTable— process-wide intern table mappingGlobal<Value>touint32_tlabel_id with refcount. One reference per sample; the table releases theGlobalwhen the last sample referencing it is collected.HeapProfiler::SetHeapProfileSampleLabelsKey— the on/off switch and the CPED Map lookup key. When unset, zero cost in the allocation path.HeapProfiler::InternLabelValue/ReleaseLabelValue/ResolveLabelValue— embedder API for allocation trackers that need to share the intern table.No callback registration. The design associates any JS value (not just string pairs) with each sample and lets the profile consumer interpret it. The same mechanism is reusable for CPU profile samples (Szegedi's v8-dev proposal targets that use-case).
Attila Szegedi proposed a similar label mechanism for CPU profiling on v8-dev (July 2025). V8 team (Leszek Swirski) indicated they would review non-invasive patches behind
#ifdefs.Node.js changes
v8.withHeapProfileLabels(labels, fn)— runs fn with labels propagating acrossawaitv8.setHeapProfileLabels(labels)— sets labels for the current async scope (for framework middleware patterns)v8.getAllocationProfile()returnssamples[].labels(frozen, shared across samples with the same label set) and per-labelexternalBytes.ProfilingArrayBufferAllocatortracks external allocations perlabel_id(one atomic null-check per allocation/free when disabled, predicted not-taken).The JS-facing API is unchanged. Label resolution (label_id to flat array to labels object) happens in the Node.js binding with a per-call cache, not in V8.
#62273 landed the SyncHeapProfileHandle API with Symbol.dispose support. The labels API proposed here is complementary.
Motivation
In multi-tenant or multi-route Node.js servers, a memory spike today tells you how much memory grew but not what caused it. With labeled heap/external memory profiling, you can answer "route /api/search accounts for 400 MB of the 1.2 GB heap" directly from production telemetry (e.g. via OTel).
This mirrors Go's
pprof.Labelscapability.Read-time performance
The label_id design moves
getAllocationProfile()from O(total samples) to O(distinct label sets): each unique label_id is resolved once and the frozen object is shared across all samples with that id. Benchmark results (5000 retained samples, n=5 calls per run):Overhead
20-run benchmark (two-server realistic HTTP workload):
Test plan