Repository navigation
docs(nexus-pdp): correct the memory and disk sizing guidance (early access) - #670
EliMoshkovich wants to merge 7 commits into
Conversation
…ly access) PER-16873. From the PER-16040 staging scale test (nexus-pdp 0.7.0-beta.26, default storage-engine settings, 2.0M / 10.8M / 20.0M facts). Memory grows with the data set and peaks during a cold start (~6.5 GiB at 20M facts), so 'memory you configure' and a flat 4 GiB are corrected. No latency or throughput figures. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
✅ Deploy Preview for permitio-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
zeevmoney
left a comment
There was a problem hiding this comment.
Changes requested — 1 HIGH, 9 MEDIUM, 5 LOW.
Blocking:
- HIGH
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:147— no memory floor; below about 2.5 GiB Nexus PDP can stop applying updates without an error
Non-blocking:
- MEDIUM
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:142— 2.0M cold-start memory isn't from the measured build; that run peaked under load - MEDIUM
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:144— 20M serving memory "~2.7 GiB" isn't in the measurements (1.7–3.2 GiB) - MEDIUM
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:146— "+0.5" in the cold-start formula has no source; rebuilds not covered - MEDIUM
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:140— peaks are sampled lower bounds; runs had 8/32 GiB limits - MEDIUM
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:150— 5.8 GiB reading doesn't show the disk guidance holds - MEDIUM
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:128— sentence above the caution still says memory depends on cache sizes you set - MEDIUM
docs/concepts/pdp/nexus-pdp.mdx:61— the memory limit, not a setting, sizes the serving cache in this build - MEDIUM
docs/concepts/pdp/nexus-pdp-deployment.mdx:25— pod spec sets a memory request but no limit - MEDIUM
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:131— configuration reference and PDP overview still give the old memory guidance - LOW
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:148— cold-start growth claim needs its CPU caveat - LOW
docs/concepts/pdp/nexus-pdp-deployment.mdx:31— name the startup-probe endpoint and the CPU count - LOW
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:152— "deliberately heavier than a typical one" has no source - LOW
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:123— "grows far more slowly" has no like-for-like measurement - LOW
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:138— small consistency fixes in the new section
Details are in the inline comments on each line.
…873) - Add the 2.5 GiB memory floor (PER-16808) to the caution, the sizing bullet, the deployment Memory row and the configuration reference. - Describe the serving cache as sized from the container memory limit (SurrealDB 3.2.4 defaults); scope the SURREAL_ROCKSDB_* variables to the snapshot load; add a memory limit to the pod spec. - 2.0M row: cold-start peak "Not recorded"; serving ranges from the 10 s container working set (2.1-2.9 GiB at 10.8M, 1.7-3.2 GiB at 20M); note the peaks are sampled lower bounds and the runs' 8/32 GiB limits. - Scope "peaks during a cold start" to large data sets; drop the unsourced +0.5 GiB end of the formula; note a rebuild was not measured. - Disk: report measured use (0.33 / 1.6 / up to 5.8 GiB) and size the volume from the snapshot (3x cold start, 4x rebuild, unbounded event store). - Cold start: CPU caveat, /health/ready startup probe, the node wait. - Cite the container PDP's 6 KB-per-object estimate as an estimate. - Consistency: Google Drive-style, role assignments in the change stream, image digest, facts link on the deployment page. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…-sizing' into eli/per-16873-nexus-pdp-measured-sizing
…ecovery (PER-16873) - 2304 MiB (2.25 GiB) is the stall point; 2.5 GiB the minimum to run at. - Nexus PDP keeps answering from the data it already has; beta.26 reports nothing, and at 2 CPUs or fewer the process can stop answering. - Restart to recover; decimal memory units count as written. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Follow-up on the two points that were open before customer-facing use. Both are now backed by a test or the owner's ticket, not only the source. 1. The memory limit sizes the serving database;
This matches the table in the configuration reference, and PER-16808's bisect ("the embedded SDK ignores 2. The stall caution now matches PER-16808 and @omer9564's draft customer page (cloud-pdp #270,
When a release carries cloud-pdp #306 (refuse to start below the floor) and #292 ( 🤖 Generated with Claude Code |
zeevmoney
left a comment
There was a problem hiding this comment.
Approved — no CRITICAL or HIGH issues found.
Non-blocking:
- MEDIUM
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:150— the recommended limits cover a first boot, not a rebuild - LOW
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:144— the 2.0M row's serving range mixes two builds - LOW
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:131— the snapshot load isn't what stalls, and the floor is worded differently across pages - LOW
docs/concepts/pdp/nexus-pdp-configuration.mdx:84— theSURREAL_ROCKSDB_*variables barely affect the snapshot load either - LOW
docs/concepts/pdp/overview.mdx:53— the PDP overview doesn't say that Nexus PDP memory still grows with the data set - LOW
docs/concepts/pdp/nexus-pdp-how-it-works.mdx:162— disk sizing assumes a known snapshot size, and a killed cold start leaves its copies behind - LOW
docs/concepts/pdp/nexus-pdp-deployment.mdx:31— the startup probe times are from 8 CPUs only
Details are in the inline comments on each line.
- Size memory for a rebuild: 4x the snapshot + 1 GiB (about 7 / 12 GiB at 10M / 20M facts), labelled estimates, on all three pages. - 2.0M serving memory: up to 1.5 GiB (the measured build only). - The stall comes from applying changes, not the SST snapshot load; the 2304 MiB floor is worded the same everywhere; raise the limit to recover. - SURREAL_ROCKSDB_* barely affect the snapshot load either. - PDP overview: memory still grows with the data set. - Disk: snapshot size per million facts; a killed cold start leaves copies. - Startup probe: fewer CPUs take longer; startupProbe in the pod spec. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rrections (PER-16873) Omer: customers want permission-check benchmarks, not cold-start figures, and no Nexus PDP numbers should be published until a proper check benchmark exists (PER-16851). Removed the measured section (table, per-size limits, startup times, snapshot growth, the 36 GB comparison) and every number derived from it. Kept the corrections to wrong public guidance: the 2.5 GiB floor, the cache sized from the memory limit, SURREAL_ROCKSDB_* scope, the pod-spec memory limit and startup probe, and the disk guidance. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@omer9564 Done in d3479ff, per your note. The scale-test measurements are gone: the cold start / memory table, the per-size limits, the startup times, the snapshot growth and the 36 GB comparison. No Nexus PDP numbers are published. What's left only corrects guidance that was wrong for the current build: the 2.5 GiB floor (your PER-16808 framing), the cache sized from the memory limit, the |
PER-16873. Corrects the Nexus PDP memory and disk guidance in the early-access docs. It publishes no performance numbers.
Why
The Nexus PDP pages said memory is "limited by a cache size you configure" and that a few hundred MiB OOMs at startup. Both are wrong for the current build, and following them can cause failures:
SURREAL_ROCKSDB_*variables don't change the serving database's memory.What
SURREAL_ROCKSDB_*only touch the snapshot-file step.limits.memoryand astartupProbeon/health/readyin the pod spec.What it doesn't publish
No performance numbers, after Omer's feedback. Customers want permission-check benchmarks, and Nexus PDP numbers should wait for a proper check benchmark like the Cloud PDP one (PER-16851). The scale-test measurements that were in earlier revisions are removed: the cold start and memory table, the per-size limits, the startup times, the snapshot growth and the 36 GB comparison. The existing "No published Nexus PDP throughput figures" caution stays.
How it was checked
b57de0a): the SST ingest, the leftover staging directory, health reporting.permitio/nexus-pdp:0.7.0-beta.26, reading the serving database's RocksDB OPTIONS at different memory limits: write buffers follow the limit, andSURREAL_ROCKSDB_*are ignored. Details in this comment.Follow-up: PER-16887 updates these pages when a release carries cloud-pdp #306, #292 or #299.
🤖 Generated with Claude Code