Skip to content

docs(nexus-pdp): correct the memory and disk sizing guidance (early access) - #670

Open
EliMoshkovich wants to merge 7 commits into
masterfrom
eli/per-16873-nexus-pdp-measured-sizing
Open

EliMoshkovich wants to merge 7 commits into
masterfrom
eli/per-16873-nexus-pdp-measured-sizing

Conversation

@EliMoshkovich

@EliMoshkovich EliMoshkovich commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

PER-16873. Corrects the Nexus PDP memory and disk guidance in the early-access docs. It publishes no performance numbers.

Why

The Nexus PDP pages said memory is "limited by a cache size you configure" and that a few hundred MiB OOMs at startup. Both are wrong for the current build, and following them can cause failures:

  • Below 2304 MiB, the embedded database can stop applying changes silently (PER-16808).
  • The documented SURREAL_ROCKSDB_* variables don't change the serving database's memory.
  • The pod-spec example sets no memory limit, so the cache is sized from the whole node.

What

  • Memory floor: never less than 2.5 GiB; the stall point is 2304 MiB. The text follows PER-16808 and Omer's draft (cloud-pdp Raz/per 9300 sdk feature parity page for docs #270): the PDP keeps answering from the data it has, this build reports nothing, and to recover you raise the limit and restart.
  • How memory is sized: the serving database sizes its cache and write buffers from the container's memory limit, or from the node's memory if there's no limit. Memory grows with the data set and peaks while a snapshot loads (cold start or rebuild). The configuration reference documents the limit-derived sizes, and says SURREAL_ROCKSDB_* only touch the snapshot-file step.
  • Deployment page: a memory-limit row, limits.memory and a startupProbe on /health/ready in the pod spec.
  • Disk: about 3× the snapshot during a cold start and at least 4× during a rebuild. A killed cold start can leave copies behind, and the event store has no size limit, so monitor free space.

What it doesn't publish

No performance numbers, after Omer's feedback. Customers want permission-check benchmarks, and Nexus PDP numbers should wait for a proper check benchmark like the Cloud PDP one (PER-16851). The scale-test measurements that were in earlier revisions are removed: the cold start and memory table, the per-size limits, the startup times, the snapshot growth and the 36 GB comparison. The existing "No published Nexus PDP throughput figures" caution stays.

How it was checked

  • The source of the measured build (cloud-pdp b57de0a): the SST ingest, the leftover staging directory, health reporting.
  • A local, offline run of permitio/nexus-pdp:0.7.0-beta.26, reading the serving database's RocksDB OPTIONS at different memory limits: write buffers follow the limit, and SURREAL_ROCKSDB_* are ignored. Details in this comment.

Follow-up: PER-16887 updates these pages when a release carries cloud-pdp #306, #292 or #299.

🤖 Generated with Claude Code

…ly access)

PER-16873. From the PER-16040 staging scale test (nexus-pdp 0.7.0-beta.26,
default storage-engine settings, 2.0M / 10.8M / 20.0M facts). Memory grows
with the data set and peaks during a cold start (~6.5 GiB at 20M facts),
so 'memory you configure' and a flat 4 GiB are corrected. No latency or
throughput figures.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@EliMoshkovich
EliMoshkovich requested review from omer9564 and zeevmoney and a balanced review from Copilot October 5, 2026 21:51
@netlify

netlify Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for permitio-docs ready!

Name Link
🔨 Latest commit d3479ff
🔍 Latest deploy log https://app.netlify.com/projects/permitio-docs/deploys/6ac5686a3738860008171c59
😎 Deploy Preview https://deploy-preview-670--permitio-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@linear-code

linear-code Bot commented Oct 5, 2026

Copy link
Copy Markdown

PER-16873

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@zeevmoney zeevmoney left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes requested — 1 HIGH, 9 MEDIUM, 5 LOW.

Blocking:

  • HIGH docs/concepts/pdp/nexus-pdp-how-it-works.mdx:147 — no memory floor; below about 2.5 GiB Nexus PDP can stop applying updates without an error

Non-blocking:

  • MEDIUM docs/concepts/pdp/nexus-pdp-how-it-works.mdx:142 — 2.0M cold-start memory isn't from the measured build; that run peaked under load
  • MEDIUM docs/concepts/pdp/nexus-pdp-how-it-works.mdx:144 — 20M serving memory "~2.7 GiB" isn't in the measurements (1.7–3.2 GiB)
  • MEDIUM docs/concepts/pdp/nexus-pdp-how-it-works.mdx:146 — "+0.5" in the cold-start formula has no source; rebuilds not covered
  • MEDIUM docs/concepts/pdp/nexus-pdp-how-it-works.mdx:140 — peaks are sampled lower bounds; runs had 8/32 GiB limits
  • MEDIUM docs/concepts/pdp/nexus-pdp-how-it-works.mdx:150 — 5.8 GiB reading doesn't show the disk guidance holds
  • MEDIUM docs/concepts/pdp/nexus-pdp-how-it-works.mdx:128 — sentence above the caution still says memory depends on cache sizes you set
  • MEDIUM docs/concepts/pdp/nexus-pdp.mdx:61 — the memory limit, not a setting, sizes the serving cache in this build
  • MEDIUM docs/concepts/pdp/nexus-pdp-deployment.mdx:25 — pod spec sets a memory request but no limit
  • MEDIUM docs/concepts/pdp/nexus-pdp-how-it-works.mdx:131 — configuration reference and PDP overview still give the old memory guidance
  • LOW docs/concepts/pdp/nexus-pdp-how-it-works.mdx:148 — cold-start growth claim needs its CPU caveat
  • LOW docs/concepts/pdp/nexus-pdp-deployment.mdx:31 — name the startup-probe endpoint and the CPU count
  • LOW docs/concepts/pdp/nexus-pdp-how-it-works.mdx:152 — "deliberately heavier than a typical one" has no source
  • LOW docs/concepts/pdp/nexus-pdp-how-it-works.mdx:123 — "grows far more slowly" has no like-for-like measurement
  • LOW docs/concepts/pdp/nexus-pdp-how-it-works.mdx:138 — small consistency fixes in the new section

Details are in the inline comments on each line.

Comment thread docs/concepts/pdp/nexus-pdp-how-it-works.mdx Outdated
Comment thread docs/concepts/pdp/nexus-pdp-how-it-works.mdx Outdated
Comment thread docs/concepts/pdp/nexus-pdp-how-it-works.mdx Outdated
Comment thread docs/concepts/pdp/nexus-pdp-how-it-works.mdx Outdated
Comment thread docs/concepts/pdp/nexus-pdp-how-it-works.mdx Outdated
Comment thread docs/concepts/pdp/nexus-pdp-how-it-works.mdx Outdated
Comment thread docs/concepts/pdp/nexus-pdp-deployment.mdx Outdated
Comment thread docs/concepts/pdp/nexus-pdp-how-it-works.mdx Outdated
Comment thread docs/concepts/pdp/nexus-pdp-how-it-works.mdx Outdated
Comment thread docs/concepts/pdp/nexus-pdp-how-it-works.mdx Outdated
Copilot AI balanced review requested due to automatic review settings October 6, 2026 17:56

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

EliMoshkovich and others added 2 commits October 6, 2026 13:06
…873)

- Add the 2.5 GiB memory floor (PER-16808) to the caution, the sizing bullet,
  the deployment Memory row and the configuration reference.
- Describe the serving cache as sized from the container memory limit
  (SurrealDB 3.2.4 defaults); scope the SURREAL_ROCKSDB_* variables to the
  snapshot load; add a memory limit to the pod spec.
- 2.0M row: cold-start peak "Not recorded"; serving ranges from the 10 s
  container working set (2.1-2.9 GiB at 10.8M, 1.7-3.2 GiB at 20M); note the
  peaks are sampled lower bounds and the runs' 8/32 GiB limits.
- Scope "peaks during a cold start" to large data sets; drop the unsourced
  +0.5 GiB end of the formula; note a rebuild was not measured.
- Disk: report measured use (0.33 / 1.6 / up to 5.8 GiB) and size the volume
  from the snapshot (3x cold start, 4x rebuild, unbounded event store).
- Cold start: CPU caveat, /health/ready startup probe, the node wait.
- Cite the container PDP's 6 KB-per-object estimate as an estimate.
- Consistency: Google Drive-style, role assignments in the change stream,
  image digest, facts link on the deployment page.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…-sizing' into eli/per-16873-nexus-pdp-measured-sizing
Copilot AI balanced review requested due to automatic review settings October 6, 2026 18:08

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

…ecovery (PER-16873)

- 2304 MiB (2.25 GiB) is the stall point; 2.5 GiB the minimum to run at.
- Nexus PDP keeps answering from the data it already has; beta.26 reports
  nothing, and at 2 CPUs or fewer the process can stop answering.
- Restart to recover; decimal memory units count as written.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@EliMoshkovich

Copy link
Copy Markdown
Contributor Author

Follow-up on the two points that were open before customer-facing use. Both are now backed by a test or the owner's ticket, not only the source.

1. The memory limit sizes the serving database; SURREAL_ROCKSDB_* doesn't. I tested the measured image locally (permitio/nexus-pdp:0.7.0-beta.26@sha256:fa7012ad…, --dev-fake-data-plane, --network none, fake key) and read the serving database's RocksDB OPTIONS file at each memory limit:

--memory write buffers
2g, 3g 2 × 64 MiB
4g, 8g 4 × 64 MiB
none (35 GiB Docker VM) 8 × 128 MiB
4g + SURREAL_ROCKSDB_WRITE_BUFFER_SIZE=32 MiB, …_MAX_WRITE_BUFFER_NUMBER=3, …_BLOCK_CACHE_SIZE=256 MiB 4 × 64 MiB (variables ignored)

This matches the table in the configuration reference, and PER-16808's bisect ("the embedded SDK ignores SURREAL_ROCKSDB_*"). The block cache size isn't in OPTIONS; its formula comes from SurrealDB 3.2.4's cnf.rs, as quoted in PER-16808.

2. The stall caution now matches PER-16808 and @omer9564's draft customer page (cloud-pdp #270, edge-memory.md):

  • 2304 MiB (2.25 GiB) is the stall point, and 2.5 GiB is the minimum to run at.
  • Nexus PDP keeps answering from the data it already has.
  • As of 0.7.0-beta.26 it reports nothing. At 2 CPUs or fewer the process can stop answering, which a /health liveness probe catches.
  • Restart to recover.
  • Decimal memory units count as written.

When a release carries cloud-pdp #306 (refuse to start below the floor) and #292 (ingest down, stalled_since, the write-stall metric), this caution should switch to that behavior and to the draft's "Watching for a stalled database" section. Likewise, the event-store note should change when a release carries PER-16705 (#299).

🤖 Generated with Claude Code

Copilot AI balanced review requested due to automatic review settings October 6, 2026 18:27

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@zeevmoney zeevmoney left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved — no CRITICAL or HIGH issues found.

Non-blocking:

  • MEDIUM docs/concepts/pdp/nexus-pdp-how-it-works.mdx:150 — the recommended limits cover a first boot, not a rebuild
  • LOW docs/concepts/pdp/nexus-pdp-how-it-works.mdx:144 — the 2.0M row's serving range mixes two builds
  • LOW docs/concepts/pdp/nexus-pdp-how-it-works.mdx:131 — the snapshot load isn't what stalls, and the floor is worded differently across pages
  • LOW docs/concepts/pdp/nexus-pdp-configuration.mdx:84 — the SURREAL_ROCKSDB_* variables barely affect the snapshot load either
  • LOW docs/concepts/pdp/overview.mdx:53 — the PDP overview doesn't say that Nexus PDP memory still grows with the data set
  • LOW docs/concepts/pdp/nexus-pdp-how-it-works.mdx:162 — disk sizing assumes a known snapshot size, and a killed cold start leaves its copies behind
  • LOW docs/concepts/pdp/nexus-pdp-deployment.mdx:31 — the startup probe times are from 8 CPUs only

Details are in the inline comments on each line.

Comment thread docs/concepts/pdp/nexus-pdp-how-it-works.mdx Outdated
Comment thread docs/concepts/pdp/nexus-pdp-how-it-works.mdx Outdated
Comment thread docs/concepts/pdp/nexus-pdp-how-it-works.mdx Outdated
Comment thread docs/concepts/pdp/nexus-pdp-configuration.mdx Outdated
Comment thread docs/concepts/pdp/overview.mdx Outdated
Comment thread docs/concepts/pdp/nexus-pdp-how-it-works.mdx Outdated
Comment thread docs/concepts/pdp/nexus-pdp-deployment.mdx Outdated
- Size memory for a rebuild: 4x the snapshot + 1 GiB (about 7 / 12 GiB at
  10M / 20M facts), labelled estimates, on all three pages.
- 2.0M serving memory: up to 1.5 GiB (the measured build only).
- The stall comes from applying changes, not the SST snapshot load; the
  2304 MiB floor is worded the same everywhere; raise the limit to recover.
- SURREAL_ROCKSDB_* barely affect the snapshot load either.
- PDP overview: memory still grows with the data set.
- Disk: snapshot size per million facts; a killed cold start leaves copies.
- Startup probe: fewer CPUs take longer; startupProbe in the pod spec.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings October 6, 2026 18:42

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

…rrections (PER-16873)

Omer: customers want permission-check benchmarks, not cold-start figures,
and no Nexus PDP numbers should be published until a proper check benchmark
exists (PER-16851). Removed the measured section (table, per-size limits,
startup times, snapshot growth, the 36 GB comparison) and every number
derived from it. Kept the corrections to wrong public guidance: the 2.5 GiB
floor, the cache sized from the memory limit, SURREAL_ROCKSDB_* scope, the
pod-spec memory limit and startup probe, and the disk guidance.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings October 6, 2026 21:30
@EliMoshkovich EliMoshkovich changed the title docs(nexus-pdp): measured memory and cold start by data-set size (early access) docs(nexus-pdp): correct the memory and disk sizing guidance (early access) Oct 6, 2026
@EliMoshkovich

Copy link
Copy Markdown
Contributor Author

@omer9564 Done in d3479ff, per your note. The scale-test measurements are gone: the cold start / memory table, the per-size limits, the startup times, the snapshot growth and the 36 GB comparison. No Nexus PDP numbers are published. What's left only corrects guidance that was wrong for the current build: the 2.5 GiB floor (your PER-16808 framing), the cache sized from the memory limit, the SURREAL_ROCKSDB_* scope, the pod-spec memory limit and startup probe, and the disk guidance. Check-latency numbers wait for PER-16851. Could you take a look?

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants