From e18e8e5d17fd0b4209a58a2f4bbd09cb7e54201e Mon Sep 17 00:00:00 2001 From: eli Date: Mon, 5 Oct 2026 16:51:05 -0500 Subject: [PATCH 1/5] docs(nexus-pdp): measured memory and cold start by data-set size (early access) PER-16873. From the PER-16040 staging scale test (nexus-pdp 0.7.0-beta.26, default storage-engine settings, 2.0M / 10.8M / 20.0M facts). Memory grows with the data set and peaks during a cold start (~6.5 GiB at 20M facts), so 'memory you configure' and a flat 4 GiB are corrected. No latency or throughput figures. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/concepts/pdp/nexus-pdp-deployment.mdx | 4 ++-- docs/concepts/pdp/nexus-pdp-how-it-works.mdx | 22 ++++++++++++++++++-- docs/concepts/pdp/nexus-pdp.mdx | 4 ++-- 3 files changed, 24 insertions(+), 6 deletions(-) diff --git a/docs/concepts/pdp/nexus-pdp-deployment.mdx b/docs/concepts/pdp/nexus-pdp-deployment.mdx index 809a52a4..b1709d2f 100644 --- a/docs/concepts/pdp/nexus-pdp-deployment.mdx +++ b/docs/concepts/pdp/nexus-pdp-deployment.mdx @@ -22,13 +22,13 @@ Nexus PDP has operational requirements that the container PDP (the Edge PDP imag | Image | `permitio/nexus-pdp`, pinned by digest (`permitio/nexus-pdp@sha256:`) | `permitio/nexus-pdp` has no `latest` tag, so a pull without a tag fails. Tags can move: `0-beta` moves to each new beta or release build, and a numbered beta tag name can be published again for a later build. A new build can rename configuration variables. See [Nexus PDP configuration reference](/concepts/pdp/nexus-pdp-configuration). | | One container per environment | Set `PDP_API_KEY` to the Nexus PDP API key of one Permit environment | Nexus PDP has no multi-environment mode. The Nexus PDP API key binds the container to exactly one environment. | | Persistent storage | Mount a persistent volume at `/var/lib/edge-pdp`, which holds the embedded database and the event store at default paths | On ephemeral storage, every restart runs a full cold start with a snapshot transfer of your whole data set. | -| Memory | 4 GiB to start | At default storage-engine settings, a container with a few hundred MiB is killed for running out of memory (OOM) at startup. See [Nexus PDP resource footprint](/concepts/pdp/nexus-pdp-how-it-works#resource-footprint). | +| Memory | 4 GiB to start; about 6 GiB at 10 million facts and 10 GiB at 20 million | At default storage-engine settings, a container with a few hundred MiB is killed for running out of memory (OOM) at startup, and a large data set needs more memory during a cold start than afterward. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint). | | Volume permissions | The volume is writable by user ID and group ID `10001` | Nexus PDP runs as the non-root user `10001` and cannot write its database or event store. | | Exposed ports | `7000` (authorization API) and `7001` (health) only | Other ports bind to loopback inside the container and are not reachable from outside. | | Termination grace period | `terminationGracePeriodSeconds: 40` or higher | 40 seconds is the sum of the default shutdown budgets `EDGE_DRAIN_TIMEOUT_SECS` (10) and `EDGE_CHILD_TERMINATION_TIMEOUT_SECS` (30). With a shorter grace period, Kubernetes kills the NATS leaf node before its on-disk event store finishes flushing. | | Liveness probe | `GET /health` on port `7001` | A liveness probe on port `7000` fails during the cold start and restarts the container. See [Health and readiness](#health-and-readiness). | | Readiness probe | `GET /health/ready` on port `7001` | Traffic reaches a Nexus PDP that cannot answer yet. | -| Startup probe | Enough time for a cold start of your data set | A cold start takes longer as your data grows. A short startup probe restarts a first boot that is still loading data. | +| Startup probe | Enough time for a cold start of your data set, with headroom (measured: 68 s at 10.8 million facts, 151 s at 20 million) | A cold start takes longer as your data grows. A short startup probe restarts a first boot that is still loading data. | ## Run Nexus PDP with Docker diff --git a/docs/concepts/pdp/nexus-pdp-how-it-works.mdx b/docs/concepts/pdp/nexus-pdp-how-it-works.mdx index 32e9d0db..456ae869 100644 --- a/docs/concepts/pdp/nexus-pdp-how-it-works.mdx +++ b/docs/concepts/pdp/nexus-pdp-how-it-works.mdx @@ -120,7 +120,7 @@ Nexus PDP stores authorization data in a different place than the container PDP, | | Container PDP (`pdp-v2`) | Nexus PDP (`nexus-pdp`) | | --- | --- | --- | | Authorization data | In OPA's in-memory document | On disk, in an embedded database | -| Memory as data grows | Grows with your data set | Limited by a cache size you configure | +| Memory as data grows | Grows with your data set | Grows far more slowly, and peaks during a cold start | | Processes | Rust API server, Python OPAL client (Horizon), and OPA | Rust binary, NATS leaf node, and OPA | | Python runtime | Required, for the Open Policy Administration Layer (OPAL) client | Not present | | Persistent storage | Not required | Required | @@ -128,11 +128,29 @@ Nexus PDP stores authorization data in a different place than the container PDP, On the container PDP, your authorization data is raw JSON in memory. The [OPA policy performance documentation](https://www.openpolicyagent.org/docs/policy-performance) states that raw JSON in OPA uses about 20 times the memory of the same data in a compact on-disk form. On Nexus PDP, the data set lives on disk in a compact form, and memory depends on the storage engine's cache sizes, which you set. :::caution Start Nexus PDP with 4 GiB of memory -The default storage-engine settings of Nexus PDP target large workloads: a 512 MiB block cache and 256 MiB write buffers. If you give the container a memory limit of a few hundred MiB, the container is killed for running out of memory (OOM) at startup. Give the Nexus PDP container 4 GiB of memory to start. +The default storage-engine settings of Nexus PDP target large workloads: a 512 MiB block cache and 256 MiB write buffers. If you give the container a memory limit of a few hundred MiB, the container is killed for running out of memory (OOM) at startup. Give the Nexus PDP container 4 GiB of memory to start, and more for a large data set: see [Measured memory and cold start](#measured-footprint). To run Nexus PDP with less memory, lower the cache and write-buffer sizes in [Nexus PDP storage engine settings](/concepts/pdp/nexus-pdp-configuration#storage-engine), then load test against your own data set before you choose a size. ::: +### Measured memory and cold start (early access) \{#measured-footprint} + +Permit measured these figures in October 2026 on Nexus PDP `0.7.0-beta.26`, an early-access build, with default storage-engine settings. The data set was one Google Drive–style ReBAC model (tenants, groups, folders, and documents, nested up to five levels), loaded in three steps. Facts are the rows Permit stores for an environment: users, tenants, resource instances, role assignments, and relationship tuples. Treat the figures as a starting point, not a guarantee: later builds can change them, and your policy and data shape change them too. + +| Facts | Users / resource instances | CPUs | Cold start, to correct data | Peak memory, during the cold start | Memory while serving checks | Snapshot size | +| --- | --- | --- | --- | --- | --- | --- | +| 2.0M | 182K / 402K | 2 | 24 s | ~0.6 GiB | 1.0–1.5 GiB | 0.26 GiB | +| 10.8M | 1.0M / 2.2M | 8 | 68 s | ~3.9 GiB | 2.1–3.0 GiB | 1.35 GiB | +| 20.0M | 1.85M / 4.1M | 8 | 151 s | ~6.5 GiB | ~2.7 GiB | 2.67 GiB | + +- **Memory peaks during a cold start.** Loading the snapshot takes about twice the snapshot's size plus 0.5–1.2 GiB. Size the memory limit for that peak, not for steady state, or the container is killed while it loads. +- **Memory limit by data-set size.** 4 GiB covers a few million facts. Give 6 GiB at about 10 million facts and 10 GiB at about 20 million. Permit has not measured larger data sets: load test before you choose a size. +- **Cold start time grows faster than the data set.** Set the startup probe from the cold start time for your size, with headroom. A cold start on a new Kubernetes node also waits for the node and the image pull. +- **The snapshot grows linearly**, at about 0.13 GiB for each million facts. +- **Disk.** At 20 million facts, the volume held 5.8 GiB, about twice the snapshot, which matches the [disk guidance](#disk). + +These figures cover memory, disk, and cold start only. They are not latency or throughput figures: the test's data set was deliberately heavier than a typical one, and the [caution on throughput](#running-at-high-volume) still applies. + ### Disk Two Nexus PDP paths must be on persistent storage: the embedded database (`EDGE_DB_PATH`) and the event store (`EDGE_DATA_DIR`). Plan for about twice the size of your data set, plus headroom. During a rebuild, the old and new database copies exist on disk together until Nexus PDP swaps in the new copy. diff --git a/docs/concepts/pdp/nexus-pdp.mdx b/docs/concepts/pdp/nexus-pdp.mdx index a0eaa0c4..09fb81b6 100644 --- a/docs/concepts/pdp/nexus-pdp.mdx +++ b/docs/concepts/pdp/nexus-pdp.mdx @@ -47,7 +47,7 @@ Nexus PDP stores relationship data in an embedded database built for graph trave The container PDP holds all authorization data in memory, so a large environment needs a large container PDP. The [OPA policy performance documentation](https://www.openpolicyagent.org/docs/policy-performance) states that raw JSON loaded into OPA uses about 20 times the memory of the same data in a compact serialized form. -Nexus PDP stores the compact form on disk and keeps a cache of configurable size in memory. You set the memory budget instead of deriving it from the size of your data. +Nexus PDP stores the compact form on disk and keeps a cache of configurable size in memory. Its memory still grows with your data set, but far more slowly, and peaks during a cold start. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint) for figures by data-set size. ### Container PDP updates need a second request \{#updates-cost-a-round-trip} @@ -58,7 +58,7 @@ Nexus PDP receives each update with the changed data inside the message. The con ## What Nexus PDP provides \{#what-it-gives-you} - **Decisions inside the container.** Every hop in a decision runs over loopback or reads local disk. No authorization query depends on reaching Permit. -- **Memory you configure.** The data set lives on disk, and a configurable cache limits resident memory. +- **Less memory for the same data.** The data set lives on disk, and a configurable cache holds the hot part of it in memory. - **Sync that resumes after a disconnection.** The control plane keeps changes for each Nexus PDP until that PDP acknowledges them, so a reconnecting Nexus PDP does not fetch its data set again. - **Decisions during control-plane outages.** When the control plane is unreachable, Nexus PDP keeps answering from its local copy and reports how stale that copy is on its health endpoint. From 8e03ba2bafcd63fb3fc03ecfe1b159a6f32a9896 Mon Sep 17 00:00:00 2001 From: eli Date: Tue, 6 Oct 2026 13:06:55 -0500 Subject: [PATCH 2/5] docs(nexus-pdp): address Zeev's review of the measured sizing (PER-16873) - Add the 2.5 GiB memory floor (PER-16808) to the caution, the sizing bullet, the deployment Memory row and the configuration reference. - Describe the serving cache as sized from the container memory limit (SurrealDB 3.2.4 defaults); scope the SURREAL_ROCKSDB_* variables to the snapshot load; add a memory limit to the pod spec. - 2.0M row: cold-start peak "Not recorded"; serving ranges from the 10 s container working set (2.1-2.9 GiB at 10.8M, 1.7-3.2 GiB at 20M); note the peaks are sampled lower bounds and the runs' 8/32 GiB limits. - Scope "peaks during a cold start" to large data sets; drop the unsourced +0.5 GiB end of the formula; note a rebuild was not measured. - Disk: report measured use (0.33 / 1.6 / up to 5.8 GiB) and size the volume from the snapshot (3x cold start, 4x rebuild, unbounded event store). - Cold start: CPU caveat, /health/ready startup probe, the node wait. - Cite the container PDP's 6 KB-per-object estimate as an estimate. - Consistency: Google Drive-style, role assignments in the change stream, image digest, facts link on the deployment page. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/concepts/pdp/nexus-pdp-configuration.mdx | 18 ++++++-- docs/concepts/pdp/nexus-pdp-deployment.mdx | 8 ++-- docs/concepts/pdp/nexus-pdp-how-it-works.mdx | 42 +++++++++++-------- docs/concepts/pdp/nexus-pdp.mdx | 4 +- docs/concepts/pdp/overview.mdx | 2 +- 5 files changed, 47 insertions(+), 27 deletions(-) diff --git a/docs/concepts/pdp/nexus-pdp-configuration.mdx b/docs/concepts/pdp/nexus-pdp-configuration.mdx index 1eb81925..6edd408d 100644 --- a/docs/concepts/pdp/nexus-pdp-configuration.mdx +++ b/docs/concepts/pdp/nexus-pdp-configuration.mdx @@ -71,16 +71,26 @@ Set the Kubernetes `terminationGracePeriodSeconds` to at least `EDGE_DRAIN_TIMEO ## Storage engine -The Nexus PDP embedded database (SurrealDB on RocksDB) uses storage-engine defaults sized for large workloads, not for a small container. These variables have the largest effect on memory use: +The Nexus PDP embedded database is SurrealDB on RocksDB. As of `0.7.0-beta.26`, the database that serves checks sizes its memory from the container's memory limit, not from environment variables: -| Variable | Default | Effect | +| Setting | Size | +| --- | --- | +| Block cache (the read cache) | Half the memory limit minus 1 GiB, and at least 16 MiB | +| Size of each write buffer | 32 MiB below a 1 GiB limit, 64 MiB below 16 GiB, and 128 MiB from 16 GiB | +| Write buffers in memory at once | 2 below a 4 GiB limit, 4 below 16 GiB, 8 below 64 GiB, and 32 from 64 GiB | + +If the container has no memory limit, the database sizes these from the node's total memory instead. Set a memory limit, for example `resources.limits.memory` on Kubernetes. + +These variables apply only to the step of a cold start that loads the snapshot files into a new database. They don't change the memory of the database that serves checks: + +| Variable | Default | Effect during the snapshot load | | --- | --- | --- | -| `SURREAL_ROCKSDB_BLOCK_CACHE_SIZE` | `536870912` (512 MiB) | Size of the read cache. The largest contributor to resident memory. | +| `SURREAL_ROCKSDB_BLOCK_CACHE_SIZE` | `536870912` (512 MiB) | Size of the read cache. | | `SURREAL_ROCKSDB_WRITE_BUFFER_SIZE` | `268435456` (256 MiB) | Size of each in-memory write buffer before RocksDB flushes it to disk. | | `SURREAL_ROCKSDB_MAX_WRITE_BUFFER_NUMBER` | `32` | Maximum number of write buffers in memory at once. | | `SURREAL_ROCKSDB_BACKGROUND_THREADS` | `4` | Number of background compaction threads. | -At these defaults, a container with a memory limit of a few hundred MiB is killed for running out of memory (OOM) at startup. Start with 4 GiB of memory, then lower these values while you load test against your own data set. See [Nexus PDP resource footprint](/concepts/pdp/nexus-pdp-how-it-works#resource-footprint). +Give Nexus PDP 4 GiB of memory for a few million facts and more for a larger data set, and never less than 2.5 GiB. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint). ## Next steps \{#related-documentation} diff --git a/docs/concepts/pdp/nexus-pdp-deployment.mdx b/docs/concepts/pdp/nexus-pdp-deployment.mdx index b1709d2f..3e8d49a8 100644 --- a/docs/concepts/pdp/nexus-pdp-deployment.mdx +++ b/docs/concepts/pdp/nexus-pdp-deployment.mdx @@ -22,13 +22,13 @@ Nexus PDP has operational requirements that the container PDP (the Edge PDP imag | Image | `permitio/nexus-pdp`, pinned by digest (`permitio/nexus-pdp@sha256:`) | `permitio/nexus-pdp` has no `latest` tag, so a pull without a tag fails. Tags can move: `0-beta` moves to each new beta or release build, and a numbered beta tag name can be published again for a later build. A new build can rename configuration variables. See [Nexus PDP configuration reference](/concepts/pdp/nexus-pdp-configuration). | | One container per environment | Set `PDP_API_KEY` to the Nexus PDP API key of one Permit environment | Nexus PDP has no multi-environment mode. The Nexus PDP API key binds the container to exactly one environment. | | Persistent storage | Mount a persistent volume at `/var/lib/edge-pdp`, which holds the embedded database and the event store at default paths | On ephemeral storage, every restart runs a full cold start with a snapshot transfer of your whole data set. | -| Memory | 4 GiB to start; about 6 GiB at 10 million facts and 10 GiB at 20 million | At default storage-engine settings, a container with a few hundred MiB is killed for running out of memory (OOM) at startup, and a large data set needs more memory during a cold start than afterward. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint). | +| Memory limit (`resources.limits.memory`) | 4 GiB to start, and never less than 2.5 GiB; about 6 GiB at 10 million [facts](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint) and 10 GiB at 20 million | Below about 2.25 GiB, the embedded database can stop applying updates without reporting an error. A large data set needs more memory during a cold start than afterward, and the container is killed if the cold start passes the limit. The embedded database sizes its cache from the memory limit; without a limit, it sizes the cache from the node's memory. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint). | | Volume permissions | The volume is writable by user ID and group ID `10001` | Nexus PDP runs as the non-root user `10001` and cannot write its database or event store. | | Exposed ports | `7000` (authorization API) and `7001` (health) only | Other ports bind to loopback inside the container and are not reachable from outside. | | Termination grace period | `terminationGracePeriodSeconds: 40` or higher | 40 seconds is the sum of the default shutdown budgets `EDGE_DRAIN_TIMEOUT_SECS` (10) and `EDGE_CHILD_TERMINATION_TIMEOUT_SECS` (30). With a shorter grace period, Kubernetes kills the NATS leaf node before its on-disk event store finishes flushing. | | Liveness probe | `GET /health` on port `7001` | A liveness probe on port `7000` fails during the cold start and restarts the container. See [Health and readiness](#health-and-readiness). | | Readiness probe | `GET /health/ready` on port `7001` | Traffic reaches a Nexus PDP that cannot answer yet. | -| Startup probe | Enough time for a cold start of your data set, with headroom (measured: 68 s at 10.8 million facts, 151 s at 20 million) | A cold start takes longer as your data grows. A short startup probe restarts a first boot that is still loading data. | +| Startup probe | `GET /health/ready` on port `7001`, with `failureThreshold` × `periodSeconds` longer than your cold start plus headroom (measured on 8 CPUs: 68 s at 10.8 million facts, 151 s at 20 million) | A cold start takes longer as your data grows. A short startup probe restarts a first boot that is still loading data. | ## Run Nexus PDP with Docker @@ -45,7 +45,7 @@ docker run -d --name nexus-pdp \ ## Kubernetes pod settings for Nexus PDP -This pod-spec excerpt sets the ports, probes, memory request, volume, and shutdown budget from the requirements table. `fsGroup: 10001` makes the mounted volume writable by the non-root user that Nexus PDP runs as. Create the `nexus-pdp-api-key` secret with the environment's Nexus PDP API key under the key `PDP_API_KEY`, which the pod spec reads, reusing the shell variable from the Docker step: +This pod-spec excerpt sets the ports, probes, memory request and limit, volume, and shutdown budget from the requirements table. `fsGroup: 10001` makes the mounted volume writable by the non-root user that Nexus PDP runs as. Create the `nexus-pdp-api-key` secret with the environment's Nexus PDP API key under the key `PDP_API_KEY`, which the pod spec reads, reusing the shell variable from the Docker step: ```bash kubectl create secret generic nexus-pdp-api-key \ @@ -67,6 +67,8 @@ containers: resources: requests: memory: 4Gi + limits: + memory: 4Gi env: - name: PDP_API_KEY valueFrom: diff --git a/docs/concepts/pdp/nexus-pdp-how-it-works.mdx b/docs/concepts/pdp/nexus-pdp-how-it-works.mdx index 456ae869..df51d15e 100644 --- a/docs/concepts/pdp/nexus-pdp-how-it-works.mdx +++ b/docs/concepts/pdp/nexus-pdp-how-it-works.mdx @@ -15,7 +15,7 @@ The NATS leaf node inside the Nexus PDP container keeps a persistent connection | Plane | Carries | | --- | --- | -| **Change stream** | Transactions of authorization data: users, tenants, resource instances, and relationship tuples | +| **Change stream** | Transactions of authorization data: users, tenants, resource instances, role assignments, and relationship tuples | | **Policy files** | The compiled Rego bundle that Open Policy Agent (OPA) evaluates | | **Policy schema** | Role and permission definitions from your policy | | **Snapshot** | A one-time bulk transfer of the whole data set, used only during a cold start | @@ -120,40 +120,48 @@ Nexus PDP stores authorization data in a different place than the container PDP, | | Container PDP (`pdp-v2`) | Nexus PDP (`nexus-pdp`) | | --- | --- | --- | | Authorization data | In OPA's in-memory document | On disk, in an embedded database | -| Memory as data grows | Grows with your data set | Grows far more slowly, and peaks during a cold start | +| Memory as data grows | Grows with your data set | Grows with your data set, and peaks during the cold start of a large data set | | Processes | Rust API server, Python OPAL client (Horizon), and OPA | Rust binary, NATS leaf node, and OPA | | Python runtime | Required, for the Open Policy Administration Layer (OPAL) client | Not present | | Persistent storage | Not required | Required | -On the container PDP, your authorization data is raw JSON in memory. The [OPA policy performance documentation](https://www.openpolicyagent.org/docs/policy-performance) states that raw JSON in OPA uses about 20 times the memory of the same data in a compact on-disk form. On Nexus PDP, the data set lives on disk in a compact form, and memory depends on the storage engine's cache sizes, which you set. +On the container PDP, your authorization data is raw JSON in memory. The [OPA policy performance documentation](https://www.openpolicyagent.org/docs/policy-performance) states that raw JSON in OPA uses about 20 times the memory of the same data in a compact on-disk form. On Nexus PDP, the data set lives on disk in a compact form. Memory still grows with the data set, and on a large data set it peaks during the cold start, while Nexus PDP loads the snapshot: see [Measured memory and cold start](#measured-footprint). -:::caution Start Nexus PDP with 4 GiB of memory -The default storage-engine settings of Nexus PDP target large workloads: a 512 MiB block cache and 256 MiB write buffers. If you give the container a memory limit of a few hundred MiB, the container is killed for running out of memory (OOM) at startup. Give the Nexus PDP container 4 GiB of memory to start, and more for a large data set: see [Measured memory and cold start](#measured-footprint). +:::caution Never give Nexus PDP less than 2.5 GiB of memory +Below about 2.25 GiB, the embedded database can stop applying updates for good under a sustained write load, such as a snapshot load or a burst of changes. Nexus PDP reports no error, and it can keep answering from data that no longer changes, so a revoked permission can keep being granted. A load test of permission checks alone doesn't show it, because checks don't write. -To run Nexus PDP with less memory, lower the cache and write-buffer sizes in [Nexus PDP storage engine settings](/concepts/pdp/nexus-pdp-configuration#storage-engine), then load test against your own data set before you choose a size. +Give the Nexus PDP container a memory limit of 4 GiB to start, and more for a large data set: see [Measured memory and cold start](#measured-footprint). The embedded database sizes its cache and write buffers from the container's memory limit; without a limit, it sizes them from the node's memory. See [Nexus PDP storage engine settings](/concepts/pdp/nexus-pdp-configuration#storage-engine). Before you choose a size, load test against your own data set, and include a cold start and a sustained burst of updates in the test. ::: ### Measured memory and cold start (early access) \{#measured-footprint} -Permit measured these figures in October 2026 on Nexus PDP `0.7.0-beta.26`, an early-access build, with default storage-engine settings. The data set was one Google Drive–style ReBAC model (tenants, groups, folders, and documents, nested up to five levels), loaded in three steps. Facts are the rows Permit stores for an environment: users, tenants, resource instances, role assignments, and relationship tuples. Treat the figures as a starting point, not a guarantee: later builds can change them, and your policy and data shape change them too. +Permit measured these figures in October 2026 on Nexus PDP `0.7.0-beta.26` (`permitio/nexus-pdp@sha256:fa7012adc2f2981829f914f211f599bc3e127625ec764adcb0056dbe54ff1bdb`), an early-access build, with default storage-engine settings. The data set was one Google Drive-style ReBAC model (tenants, groups, folders, and documents, nested up to five levels), loaded in three steps. Facts are the rows Permit stores for an environment: users, tenants, resource instances, role assignments, and relationship tuples. Treat the figures as a starting point, not a guarantee: later builds can change them, and your policy and data shape change them too. -| Facts | Users / resource instances | CPUs | Cold start, to correct data | Peak memory, during the cold start | Memory while serving checks | Snapshot size | +| Facts | Users / resource instances | CPUs | Cold start, to correct data | Peak memory during the cold start (at least) | Memory while serving checks | Snapshot size | | --- | --- | --- | --- | --- | --- | --- | -| 2.0M | 182K / 402K | 2 | 24 s | ~0.6 GiB | 1.0–1.5 GiB | 0.26 GiB | -| 10.8M | 1.0M / 2.2M | 8 | 68 s | ~3.9 GiB | 2.1–3.0 GiB | 1.35 GiB | -| 20.0M | 1.85M / 4.1M | 8 | 151 s | ~6.5 GiB | ~2.7 GiB | 2.67 GiB | +| 2.0M | 182K / 402K | 2 | 24 s | Not recorded | 1.0–1.5 GiB | 0.26 GiB | +| 10.8M | 1.0M / 2.2M | 8 | 68 s | ~3.9 GiB | 2.1–2.9 GiB | 1.35 GiB | +| 20.0M | 1.85M / 4.1M | 8 | 151 s | ~6.5 GiB | 1.7–3.2 GiB | 2.67 GiB | -- **Memory peaks during a cold start.** Loading the snapshot takes about twice the snapshot's size plus 0.5–1.2 GiB. Size the memory limit for that peak, not for steady state, or the container is killed while it loads. -- **Memory limit by data-set size.** 4 GiB covers a few million facts. Give 6 GiB at about 10 million facts and 10 GiB at about 20 million. Permit has not measured larger data sets: load test before you choose a size. -- **Cold start time grows faster than the data set.** Set the startup probe from the cold start time for your size, with headroom. A cold start on a new Kubernetes node also waits for the node and the image pull. +Peak memory is the highest container working set sampled every 10 seconds, so the real peak can be higher. The runs had memory limits of 8 GiB (2.0M) and 32 GiB (10.8M and 20.0M). Because the embedded database sizes its cache from the memory limit, memory while serving can differ under a smaller limit. The limits recommended below add about 50% headroom to the peaks and were not tested as limits. + +- **Memory limit by data-set size.** Never set the memory limit below 2.5 GiB, whatever the size of your data set: below that, the embedded database can stop applying updates without reporting an error. 4 GiB covers a few million facts. Give 6 GiB at about 10 million facts and 10 GiB at about 20 million. Permit has not measured larger data sets: load test before you choose a size. +- **On a large data set, memory peaks during a cold start.** Loading the snapshot takes about twice the snapshot's size on top of the rest of Nexus PDP. At 10.8 million and 20 million facts, the peak was about twice the snapshot plus 1.2 GiB. Set the memory limit above that peak with headroom, or the container is killed while it loads. A [rebuild](#self-healing) runs the same load while Nexus PDP keeps serving from the old database, so it can need more memory than a first boot. Permit has not measured a rebuild. +- **Cold start time grows with the data set.** On 8 CPUs, it took 68 s at 10.8 million facts and 151 s at 20 million: 1.85 times the facts took 2.2 times as long. Set the startup probe from the cold start time for your size, with headroom. On a new Kubernetes node, the pod also waits about a minute for the node and the image pull before the container starts. That wait counts against rollout timeouts, not the startup probe. - **The snapshot grows linearly**, at about 0.13 GiB for each million facts. -- **Disk.** At 20 million facts, the volume held 5.8 GiB, about twice the snapshot, which matches the [disk guidance](#disk). +- **Disk.** After the cold start, the volume used 0.33 GiB at 2.0 million facts and 1.6 GiB at 10.8 million. At 20 million facts, it used up to 5.8 GiB during the cold start and 2.8 GiB after it. No rebuild was measured. + +For comparison, the container PDP's [memory estimate](/how-to/deploy/deploy-to-production#scaling-memory) of 6 KB per user, resource instance, and tenant gives about 36 GB for the 20 million facts data set. That is an estimate, not a measurement of the container PDP on the same data. -These figures cover memory, disk, and cold start only. They are not latency or throughput figures: the test's data set was deliberately heavier than a typical one, and the [caution on throughput](#running-at-high-volume) still applies. +These figures cover memory, disk, and cold start only. They are not latency or throughput figures: per-check latency depends on how many resources each user reaches in your policy and data, so measure it against your own data set. The [caution on throughput](#running-at-high-volume) still applies. ### Disk -Two Nexus PDP paths must be on persistent storage: the embedded database (`EDGE_DB_PATH`) and the event store (`EDGE_DATA_DIR`). Plan for about twice the size of your data set, plus headroom. During a rebuild, the old and new database copies exist on disk together until Nexus PDP swaps in the new copy. +Two Nexus PDP paths must be on persistent storage: the embedded database (`EDGE_DB_PATH`) and the event store (`EDGE_DATA_DIR`). At the default paths, both are on the volume mounted at `/var/lib/edge-pdp`. Size the volume from your snapshot size, plus headroom: + +- **During a cold start**, the volume holds the downloaded snapshot, an unpacked copy of its files, and the database built from them, all at once: about three times the snapshot. +- **During a [rebuild](#self-healing)**, the old database also stays on the volume until Nexus PDP swaps in the new copy: about four times the snapshot. Give the volume at least four times the snapshot size. +- **The event store** keeps the changes Nexus PDP receives, with no age or size limit. It grows with your update volume over time, not with the data set, so monitor free space on the volume. ## Next steps \{#related-documentation} diff --git a/docs/concepts/pdp/nexus-pdp.mdx b/docs/concepts/pdp/nexus-pdp.mdx index 09fb81b6..1cc7cbe2 100644 --- a/docs/concepts/pdp/nexus-pdp.mdx +++ b/docs/concepts/pdp/nexus-pdp.mdx @@ -47,7 +47,7 @@ Nexus PDP stores relationship data in an embedded database built for graph trave The container PDP holds all authorization data in memory, so a large environment needs a large container PDP. The [OPA policy performance documentation](https://www.openpolicyagent.org/docs/policy-performance) states that raw JSON loaded into OPA uses about 20 times the memory of the same data in a compact serialized form. -Nexus PDP stores the compact form on disk and keeps a cache of configurable size in memory. Its memory still grows with your data set, but far more slowly, and peaks during a cold start. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint) for figures by data-set size. +Nexus PDP stores the compact form on disk and keeps a cache in memory, sized from the container's memory limit. Its memory still grows with your data set, and peaks during the cold start of a large data set. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint) for figures by data-set size. ### Container PDP updates need a second request \{#updates-cost-a-round-trip} @@ -58,7 +58,7 @@ Nexus PDP receives each update with the changed data inside the message. The con ## What Nexus PDP provides \{#what-it-gives-you} - **Decisions inside the container.** Every hop in a decision runs over loopback or reads local disk. No authorization query depends on reaching Permit. -- **Less memory for the same data.** The data set lives on disk, and a configurable cache holds the hot part of it in memory. +- **Less memory for the same data.** The data set lives on disk, and a cache sized from the container's memory limit holds the hot part of it in memory. - **Sync that resumes after a disconnection.** The control plane keeps changes for each Nexus PDP until that PDP acknowledges them, so a reconnecting Nexus PDP does not fetch its data set again. - **Decisions during control-plane outages.** When the control plane is unreachable, Nexus PDP keeps answering from its local copy and reports how stale that copy is on its health endpoint. diff --git a/docs/concepts/pdp/overview.mdx b/docs/concepts/pdp/overview.mdx index e594b8f8..ab4dff6b 100644 --- a/docs/concepts/pdp/overview.mdx +++ b/docs/concepts/pdp/overview.mdx @@ -33,7 +33,7 @@ The Edge PDP bundles three components in one container: Open Policy Agent (OPA), Nexus PDP differs from the Edge PDP in four ways: -- **Data on disk, not in memory.** The Edge PDP holds your data in OPA's in-memory JSON document, so a large environment needs a large PDP. Nexus PDP stores the data in an embedded database on disk and keeps a cache of configurable size in memory. +- **Data on disk, not in memory.** The Edge PDP holds your data in OPA's in-memory JSON document, so a large environment needs a large PDP. Nexus PDP stores the data in an embedded database on disk and keeps a cache in memory, sized from the container's memory limit. - **A store built for relationship queries.** Nexus PDP keeps relationship data in a database built for graph traversal, which OPA queries over loopback while it evaluates policy. - **Sync that resumes after disconnection.** Each update contains the changed data, and the control plane keeps each Nexus PDP's unacknowledged changes until that PDP applies them. A reconnecting Nexus PDP does not fetch its data again. - **Decisions during control-plane outages.** If the control plane is unreachable, Nexus PDP keeps answering from its local copy and reports how stale the copy is. From 8c012ff3c014c7ea618360c1fecab7d947d84840 Mon Sep 17 00:00:00 2001 From: eli Date: Tue, 6 Oct 2026 13:27:22 -0500 Subject: [PATCH 3/5] docs(nexus-pdp): state the stall floor the way PER-16808 does, with recovery (PER-16873) - 2304 MiB (2.25 GiB) is the stall point; 2.5 GiB the minimum to run at. - Nexus PDP keeps answering from the data it already has; beta.26 reports nothing, and at 2 CPUs or fewer the process can stop answering. - Restart to recover; decimal memory units count as written. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/concepts/pdp/nexus-pdp-deployment.mdx | 2 +- docs/concepts/pdp/nexus-pdp-how-it-works.mdx | 4 +++- 2 files changed, 4 insertions(+), 2 deletions(-) diff --git a/docs/concepts/pdp/nexus-pdp-deployment.mdx b/docs/concepts/pdp/nexus-pdp-deployment.mdx index 3e8d49a8..d6ddb77f 100644 --- a/docs/concepts/pdp/nexus-pdp-deployment.mdx +++ b/docs/concepts/pdp/nexus-pdp-deployment.mdx @@ -22,7 +22,7 @@ Nexus PDP has operational requirements that the container PDP (the Edge PDP imag | Image | `permitio/nexus-pdp`, pinned by digest (`permitio/nexus-pdp@sha256:`) | `permitio/nexus-pdp` has no `latest` tag, so a pull without a tag fails. Tags can move: `0-beta` moves to each new beta or release build, and a numbered beta tag name can be published again for a later build. A new build can rename configuration variables. See [Nexus PDP configuration reference](/concepts/pdp/nexus-pdp-configuration). | | One container per environment | Set `PDP_API_KEY` to the Nexus PDP API key of one Permit environment | Nexus PDP has no multi-environment mode. The Nexus PDP API key binds the container to exactly one environment. | | Persistent storage | Mount a persistent volume at `/var/lib/edge-pdp`, which holds the embedded database and the event store at default paths | On ephemeral storage, every restart runs a full cold start with a snapshot transfer of your whole data set. | -| Memory limit (`resources.limits.memory`) | 4 GiB to start, and never less than 2.5 GiB; about 6 GiB at 10 million [facts](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint) and 10 GiB at 20 million | Below about 2.25 GiB, the embedded database can stop applying updates without reporting an error. A large data set needs more memory during a cold start than afterward, and the container is killed if the cold start passes the limit. The embedded database sizes its cache from the memory limit; without a limit, it sizes the cache from the node's memory. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint). | +| Memory limit (`resources.limits.memory`) | 4 GiB to start, and never less than 2.5 GiB; about 6 GiB at 10 million [facts](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint) and 10 GiB at 20 million | Below 2304 MiB (2.25 GiB), the embedded database can stop applying changes under a sustained write load without reporting it; restart the container to recover. A large data set needs more memory during a cold start than afterward, and the container is killed if the cold start passes the limit. The embedded database sizes its cache from the memory limit; without a limit, it sizes the cache from the node's memory. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint). | | Volume permissions | The volume is writable by user ID and group ID `10001` | Nexus PDP runs as the non-root user `10001` and cannot write its database or event store. | | Exposed ports | `7000` (authorization API) and `7001` (health) only | Other ports bind to loopback inside the container and are not reachable from outside. | | Termination grace period | `terminationGracePeriodSeconds: 40` or higher | 40 seconds is the sum of the default shutdown budgets `EDGE_DRAIN_TIMEOUT_SECS` (10) and `EDGE_CHILD_TERMINATION_TIMEOUT_SECS` (30). With a shorter grace period, Kubernetes kills the NATS leaf node before its on-disk event store finishes flushing. | diff --git a/docs/concepts/pdp/nexus-pdp-how-it-works.mdx b/docs/concepts/pdp/nexus-pdp-how-it-works.mdx index df51d15e..79d86b47 100644 --- a/docs/concepts/pdp/nexus-pdp-how-it-works.mdx +++ b/docs/concepts/pdp/nexus-pdp-how-it-works.mdx @@ -128,7 +128,9 @@ Nexus PDP stores authorization data in a different place than the container PDP, On the container PDP, your authorization data is raw JSON in memory. The [OPA policy performance documentation](https://www.openpolicyagent.org/docs/policy-performance) states that raw JSON in OPA uses about 20 times the memory of the same data in a compact on-disk form. On Nexus PDP, the data set lives on disk in a compact form. Memory still grows with the data set, and on a large data set it peaks during the cold start, while Nexus PDP loads the snapshot: see [Measured memory and cold start](#measured-footprint). :::caution Never give Nexus PDP less than 2.5 GiB of memory -Below about 2.25 GiB, the embedded database can stop applying updates for good under a sustained write load, such as a snapshot load or a burst of changes. Nexus PDP reports no error, and it can keep answering from data that no longer changes, so a revoked permission can keep being granted. A load test of permission checks alone doesn't show it, because checks don't write. +Below 2304 MiB (2.25 GiB), the embedded database can stop applying changes for good under a sustained write load, such as a snapshot load or a burst of changes. Nexus PDP then keeps answering from the data it already has, so later changes to your policy data don't take effect. As of `0.7.0-beta.26`, Nexus PDP doesn't report this: `/health` stays up and nothing is logged. With 2 CPUs or fewer, the whole process can stop answering instead, and a liveness probe on `/health` restarts it. To recover, restart the container and raise its memory limit. A load test of permission checks alone doesn't show the problem, because checks don't write. + +2304 MiB is the point below which writes can stall, not a size to run at. Memory limits in decimal units count as written: `2400M` (2.4 billion bytes) is below it. Give the Nexus PDP container a memory limit of 4 GiB to start, and more for a large data set: see [Measured memory and cold start](#measured-footprint). The embedded database sizes its cache and write buffers from the container's memory limit; without a limit, it sizes them from the node's memory. See [Nexus PDP storage engine settings](/concepts/pdp/nexus-pdp-configuration#storage-engine). Before you choose a size, load test against your own data set, and include a cold start and a sustained burst of updates in the test. ::: From a4838b46d9abff35d52698119e530185224c7491 Mon Sep 17 00:00:00 2001 From: eli Date: Tue, 6 Oct 2026 13:42:27 -0500 Subject: [PATCH 4/5] docs(nexus-pdp): address Zeev's second review (PER-16873) - Size memory for a rebuild: 4x the snapshot + 1 GiB (about 7 / 12 GiB at 10M / 20M facts), labelled estimates, on all three pages. - 2.0M serving memory: up to 1.5 GiB (the measured build only). - The stall comes from applying changes, not the SST snapshot load; the 2304 MiB floor is worded the same everywhere; raise the limit to recover. - SURREAL_ROCKSDB_* barely affect the snapshot load either. - PDP overview: memory still grows with the data set. - Disk: snapshot size per million facts; a killed cold start leaves copies. - Startup probe: fewer CPUs take longer; startupProbe in the pod spec. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/concepts/pdp/nexus-pdp-configuration.mdx | 6 +++--- docs/concepts/pdp/nexus-pdp-deployment.mdx | 10 +++++++--- docs/concepts/pdp/nexus-pdp-how-it-works.mdx | 14 +++++++------- docs/concepts/pdp/overview.mdx | 2 +- 4 files changed, 18 insertions(+), 14 deletions(-) diff --git a/docs/concepts/pdp/nexus-pdp-configuration.mdx b/docs/concepts/pdp/nexus-pdp-configuration.mdx index 6edd408d..1dabf9f9 100644 --- a/docs/concepts/pdp/nexus-pdp-configuration.mdx +++ b/docs/concepts/pdp/nexus-pdp-configuration.mdx @@ -81,16 +81,16 @@ The Nexus PDP embedded database is SurrealDB on RocksDB. As of `0.7.0-beta.26`, If the container has no memory limit, the database sizes these from the node's total memory instead. Set a memory limit, for example `resources.limits.memory` on Kubernetes. -These variables apply only to the step of a cold start that loads the snapshot files into a new database. They don't change the memory of the database that serves checks: +These variables apply only to the step of a cold start or a rebuild that adds the snapshot files to a new database, and they have little effect on memory even there: that step adds whole files instead of writing through the write buffers, and its memory peak comes from holding the snapshot. They don't change the memory of the database that serves checks: -| Variable | Default | Effect during the snapshot load | +| Variable | Default | Setting it controls | | --- | --- | --- | | `SURREAL_ROCKSDB_BLOCK_CACHE_SIZE` | `536870912` (512 MiB) | Size of the read cache. | | `SURREAL_ROCKSDB_WRITE_BUFFER_SIZE` | `268435456` (256 MiB) | Size of each in-memory write buffer before RocksDB flushes it to disk. | | `SURREAL_ROCKSDB_MAX_WRITE_BUFFER_NUMBER` | `32` | Maximum number of write buffers in memory at once. | | `SURREAL_ROCKSDB_BACKGROUND_THREADS` | `4` | Number of background compaction threads. | -Give Nexus PDP 4 GiB of memory for a few million facts and more for a larger data set, and never less than 2.5 GiB. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint). +Give the Nexus PDP container a memory limit of 4 GiB for a few million facts and more for a larger data set. Never set the limit below 2.5 GiB: below 2304 MiB, the embedded database can stop applying changes without reporting it. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint). ## Next steps \{#related-documentation} diff --git a/docs/concepts/pdp/nexus-pdp-deployment.mdx b/docs/concepts/pdp/nexus-pdp-deployment.mdx index d6ddb77f..d4422500 100644 --- a/docs/concepts/pdp/nexus-pdp-deployment.mdx +++ b/docs/concepts/pdp/nexus-pdp-deployment.mdx @@ -22,13 +22,13 @@ Nexus PDP has operational requirements that the container PDP (the Edge PDP imag | Image | `permitio/nexus-pdp`, pinned by digest (`permitio/nexus-pdp@sha256:`) | `permitio/nexus-pdp` has no `latest` tag, so a pull without a tag fails. Tags can move: `0-beta` moves to each new beta or release build, and a numbered beta tag name can be published again for a later build. A new build can rename configuration variables. See [Nexus PDP configuration reference](/concepts/pdp/nexus-pdp-configuration). | | One container per environment | Set `PDP_API_KEY` to the Nexus PDP API key of one Permit environment | Nexus PDP has no multi-environment mode. The Nexus PDP API key binds the container to exactly one environment. | | Persistent storage | Mount a persistent volume at `/var/lib/edge-pdp`, which holds the embedded database and the event store at default paths | On ephemeral storage, every restart runs a full cold start with a snapshot transfer of your whole data set. | -| Memory limit (`resources.limits.memory`) | 4 GiB to start, and never less than 2.5 GiB; about 6 GiB at 10 million [facts](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint) and 10 GiB at 20 million | Below 2304 MiB (2.25 GiB), the embedded database can stop applying changes under a sustained write load without reporting it; restart the container to recover. A large data set needs more memory during a cold start than afterward, and the container is killed if the cold start passes the limit. The embedded database sizes its cache from the memory limit; without a limit, it sizes the cache from the node's memory. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint). | +| Memory limit (`resources.limits.memory`) | 4 GiB to start, and never less than 2.5 GiB; about 7 GiB at 10 million [facts](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint) and 12 GiB at 20 million (estimates that allow for a rebuild) | Below 2304 MiB (2.25 GiB), the embedded database can stop applying changes under a sustained write load without reporting it; raise the memory limit and restart the container to recover. A large data set needs more memory during a cold start than afterward, and the container is killed if the cold start passes the limit. The embedded database sizes its cache from the memory limit; without a limit, it sizes the cache from the node's memory. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint). | | Volume permissions | The volume is writable by user ID and group ID `10001` | Nexus PDP runs as the non-root user `10001` and cannot write its database or event store. | | Exposed ports | `7000` (authorization API) and `7001` (health) only | Other ports bind to loopback inside the container and are not reachable from outside. | | Termination grace period | `terminationGracePeriodSeconds: 40` or higher | 40 seconds is the sum of the default shutdown budgets `EDGE_DRAIN_TIMEOUT_SECS` (10) and `EDGE_CHILD_TERMINATION_TIMEOUT_SECS` (30). With a shorter grace period, Kubernetes kills the NATS leaf node before its on-disk event store finishes flushing. | | Liveness probe | `GET /health` on port `7001` | A liveness probe on port `7000` fails during the cold start and restarts the container. See [Health and readiness](#health-and-readiness). | | Readiness probe | `GET /health/ready` on port `7001` | Traffic reaches a Nexus PDP that cannot answer yet. | -| Startup probe | `GET /health/ready` on port `7001`, with `failureThreshold` × `periodSeconds` longer than your cold start plus headroom (measured on 8 CPUs: 68 s at 10.8 million facts, 151 s at 20 million) | A cold start takes longer as your data grows. A short startup probe restarts a first boot that is still loading data. | +| Startup probe | `GET /health/ready` on port `7001`, with `failureThreshold` × `periodSeconds` longer than your cold start plus headroom (measured on 8 CPUs: 68 s at 10.8 million facts, 151 s at 20 million). Fewer CPUs take longer, so time a cold start of your own data set before you set the probe | A cold start takes longer as your data grows. A short startup probe restarts a first boot that is still loading data. | ## Run Nexus PDP with Docker @@ -52,7 +52,7 @@ kubectl create secret generic nexus-pdp-api-key \ --from-literal=PDP_API_KEY="$NEXUS_PDP_API_KEY" ``` -Then size the `nexus-pdp-data` claim for your data set, and add a startup probe with enough time for a cold start: +Then size the `nexus-pdp-data` claim for your data set (see [Disk](/concepts/pdp/nexus-pdp-how-it-works#disk)), and add a startup probe with enough time for a cold start: ```yaml terminationGracePeriodSeconds: 40 @@ -79,6 +79,10 @@ containers: httpGet: { path: /health, port: 7001 } readinessProbe: httpGet: { path: /health/ready, port: 7001 } + startupProbe: + httpGet: { path: /health/ready, port: 7001 } + periodSeconds: 10 + failureThreshold: 30 # 300 s: set from a timed cold start of your data set volumeMounts: - name: nexus-data mountPath: /var/lib/edge-pdp diff --git a/docs/concepts/pdp/nexus-pdp-how-it-works.mdx b/docs/concepts/pdp/nexus-pdp-how-it-works.mdx index 79d86b47..302c5b32 100644 --- a/docs/concepts/pdp/nexus-pdp-how-it-works.mdx +++ b/docs/concepts/pdp/nexus-pdp-how-it-works.mdx @@ -128,7 +128,7 @@ Nexus PDP stores authorization data in a different place than the container PDP, On the container PDP, your authorization data is raw JSON in memory. The [OPA policy performance documentation](https://www.openpolicyagent.org/docs/policy-performance) states that raw JSON in OPA uses about 20 times the memory of the same data in a compact on-disk form. On Nexus PDP, the data set lives on disk in a compact form. Memory still grows with the data set, and on a large data set it peaks during the cold start, while Nexus PDP loads the snapshot: see [Measured memory and cold start](#measured-footprint). :::caution Never give Nexus PDP less than 2.5 GiB of memory -Below 2304 MiB (2.25 GiB), the embedded database can stop applying changes for good under a sustained write load, such as a snapshot load or a burst of changes. Nexus PDP then keeps answering from the data it already has, so later changes to your policy data don't take effect. As of `0.7.0-beta.26`, Nexus PDP doesn't report this: `/health` stays up and nothing is logged. With 2 CPUs or fewer, the whole process can stop answering instead, and a liveness probe on `/health` restarts it. To recover, restart the container and raise its memory limit. A load test of permission checks alone doesn't show the problem, because checks don't write. +Below 2304 MiB (2.25 GiB), the embedded database can stop applying changes for good under a sustained write load, such as applying the changes that arrived during a cold start, or a burst of changes. Nexus PDP then keeps answering from the data it already has, so later changes to your policy data don't take effect. As of `0.7.0-beta.26`, Nexus PDP doesn't report this: `/health` stays up and nothing is logged. With 2 CPUs or fewer, the whole process can stop answering instead, and a liveness probe on `/health` restarts it. To recover, restart the container and raise its memory limit. A load test of permission checks alone doesn't show the problem, because checks don't write. 2304 MiB is the point below which writes can stall, not a size to run at. Memory limits in decimal units count as written: `2400M` (2.4 billion bytes) is below it. @@ -141,14 +141,14 @@ Permit measured these figures in October 2026 on Nexus PDP `0.7.0-beta.26` (`per | Facts | Users / resource instances | CPUs | Cold start, to correct data | Peak memory during the cold start (at least) | Memory while serving checks | Snapshot size | | --- | --- | --- | --- | --- | --- | --- | -| 2.0M | 182K / 402K | 2 | 24 s | Not recorded | 1.0–1.5 GiB | 0.26 GiB | +| 2.0M | 182K / 402K | 2 | 24 s | Not recorded | Up to 1.5 GiB | 0.26 GiB | | 10.8M | 1.0M / 2.2M | 8 | 68 s | ~3.9 GiB | 2.1–2.9 GiB | 1.35 GiB | | 20.0M | 1.85M / 4.1M | 8 | 151 s | ~6.5 GiB | 1.7–3.2 GiB | 2.67 GiB | -Peak memory is the highest container working set sampled every 10 seconds, so the real peak can be higher. The runs had memory limits of 8 GiB (2.0M) and 32 GiB (10.8M and 20.0M). Because the embedded database sizes its cache from the memory limit, memory while serving can differ under a smaller limit. The limits recommended below add about 50% headroom to the peaks and were not tested as limits. +Peak memory is the highest container working set sampled every 10 seconds, so the real peak can be higher. The runs had memory limits of 8 GiB (2.0M) and 32 GiB (10.8M and 20.0M). Because the embedded database sizes its cache from the memory limit, memory while serving can differ under a smaller limit. The limits recommended below are estimates that allow for a rebuild, and were not tested as limits. -- **Memory limit by data-set size.** Never set the memory limit below 2.5 GiB, whatever the size of your data set: below that, the embedded database can stop applying updates without reporting an error. 4 GiB covers a few million facts. Give 6 GiB at about 10 million facts and 10 GiB at about 20 million. Permit has not measured larger data sets: load test before you choose a size. -- **On a large data set, memory peaks during a cold start.** Loading the snapshot takes about twice the snapshot's size on top of the rest of Nexus PDP. At 10.8 million and 20 million facts, the peak was about twice the snapshot plus 1.2 GiB. Set the memory limit above that peak with headroom, or the container is killed while it loads. A [rebuild](#self-healing) runs the same load while Nexus PDP keeps serving from the old database, so it can need more memory than a first boot. Permit has not measured a rebuild. +- **Memory limit by data-set size.** Never set the memory limit below 2.5 GiB, whatever the size of your data set: below 2304 MiB (2.25 GiB), the embedded database can stop applying changes without reporting it. 4 GiB covers a few million facts. For a larger data set, give at least four times the snapshot size plus 1 GiB, rounded up: about 7 GiB at 10 million facts and 12 GiB at 20 million. That leaves room for a [rebuild](#self-healing), which loads the snapshot while the old database keeps serving with a cache of up to half the memory limit minus 1 GiB. These limits are estimates: Permit has not measured a rebuild or larger data sets, so load test before you choose a size. +- **On a large data set, memory peaks during a cold start.** Loading the snapshot takes about twice the snapshot's size on top of the rest of Nexus PDP. At 10.8 million and 20 million facts, the peak was about twice the snapshot plus 1.2 GiB. Set the memory limit above that peak with headroom, or the container is killed while it loads. A [rebuild](#self-healing) runs the same load while Nexus PDP keeps serving from the old database, so it needs more memory than a first boot; the limits in the previous bullet allow for it. Permit has not measured a rebuild. - **Cold start time grows with the data set.** On 8 CPUs, it took 68 s at 10.8 million facts and 151 s at 20 million: 1.85 times the facts took 2.2 times as long. Set the startup probe from the cold start time for your size, with headroom. On a new Kubernetes node, the pod also waits about a minute for the node and the image pull before the container starts. That wait counts against rollout timeouts, not the startup probe. - **The snapshot grows linearly**, at about 0.13 GiB for each million facts. - **Disk.** After the cold start, the volume used 0.33 GiB at 2.0 million facts and 1.6 GiB at 10.8 million. At 20 million facts, it used up to 5.8 GiB during the cold start and 2.8 GiB after it. No rebuild was measured. @@ -159,9 +159,9 @@ These figures cover memory, disk, and cold start only. They are not latency or t ### Disk -Two Nexus PDP paths must be on persistent storage: the embedded database (`EDGE_DB_PATH`) and the event store (`EDGE_DATA_DIR`). At the default paths, both are on the volume mounted at `/var/lib/edge-pdp`. Size the volume from your snapshot size, plus headroom: +Two Nexus PDP paths must be on persistent storage: the embedded database (`EDGE_DB_PATH`) and the event store (`EDGE_DATA_DIR`). At the default paths, both are on the volume mounted at `/var/lib/edge-pdp`. Size the volume from your snapshot size, plus headroom. In the [measured data set](#measured-footprint), the snapshot was about 0.13 GiB for each million facts. -- **During a cold start**, the volume holds the downloaded snapshot, an unpacked copy of its files, and the database built from them, all at once: about three times the snapshot. +- **During a cold start**, the volume holds the downloaded snapshot, an unpacked copy of its files, and the database built from them, all at once: about three times the snapshot. If a cold start is killed before it finishes, for example by the memory limit or a short startup probe, its copies can stay on the volume, and each new attempt writes new ones. If restarts fill the volume, fix the cause, then empty the volume so the next start runs a clean cold start. - **During a [rebuild](#self-healing)**, the old database also stays on the volume until Nexus PDP swaps in the new copy: about four times the snapshot. Give the volume at least four times the snapshot size. - **The event store** keeps the changes Nexus PDP receives, with no age or size limit. It grows with your update volume over time, not with the data set, so monitor free space on the volume. diff --git a/docs/concepts/pdp/overview.mdx b/docs/concepts/pdp/overview.mdx index 7f6e8d76..564270bd 100644 --- a/docs/concepts/pdp/overview.mdx +++ b/docs/concepts/pdp/overview.mdx @@ -50,7 +50,7 @@ The Edge PDP bundles three components in one container: Open Policy Agent (OPA), Nexus PDP differs from the Edge PDP in four ways: -- **Data on disk, not in memory.** The Edge PDP holds your data in OPA's in-memory JSON document, so a large environment needs a large PDP. Nexus PDP stores the data in an embedded database on disk and keeps a cache in memory, sized from the container's memory limit. +- **Data on disk, not in memory.** The Edge PDP holds your data in OPA's in-memory JSON document, so a large environment needs a large PDP. Nexus PDP stores the data in an embedded database on disk and keeps a cache in memory, sized from the container's memory limit. Nexus PDP memory still grows with your data set, and peaks during the cold start of a large data set: see [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint). - **A store built for relationship queries.** Nexus PDP keeps relationship data in a database built for graph traversal, which OPA queries over loopback while it evaluates policy. - **Sync that resumes after disconnection.** Each update contains the changed data, and the control plane keeps each Nexus PDP's unacknowledged changes until that PDP applies them. A reconnecting Nexus PDP does not fetch its data again. - **Decisions during control-plane outages.** If the control plane is unreachable, Nexus PDP keeps answering from its local copy and reports how stale the copy is. From d3479ff34d63be4635b2cb4c1d4759592ac70b15 Mon Sep 17 00:00:00 2001 From: eli Date: Tue, 6 Oct 2026 16:30:13 -0500 Subject: [PATCH 5/5] docs(nexus-pdp): drop the scale-test measurements; keep the sizing corrections (PER-16873) Omer: customers want permission-check benchmarks, not cold-start figures, and no Nexus PDP numbers should be published until a proper check benchmark exists (PER-16851). Removed the measured section (table, per-size limits, startup times, snapshot growth, the 36 GB comparison) and every number derived from it. Kept the corrections to wrong public guidance: the 2.5 GiB floor, the cache sized from the memory limit, SURREAL_ROCKSDB_* scope, the pod-spec memory limit and startup probe, and the disk guidance. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/concepts/pdp/nexus-pdp-configuration.mdx | 2 +- docs/concepts/pdp/nexus-pdp-deployment.mdx | 4 +-- docs/concepts/pdp/nexus-pdp-how-it-works.mdx | 30 +++---------------- docs/concepts/pdp/nexus-pdp.mdx | 4 +-- docs/concepts/pdp/overview.mdx | 2 +- 5 files changed, 10 insertions(+), 32 deletions(-) diff --git a/docs/concepts/pdp/nexus-pdp-configuration.mdx b/docs/concepts/pdp/nexus-pdp-configuration.mdx index 1dabf9f9..05e9303f 100644 --- a/docs/concepts/pdp/nexus-pdp-configuration.mdx +++ b/docs/concepts/pdp/nexus-pdp-configuration.mdx @@ -90,7 +90,7 @@ These variables apply only to the step of a cold start or a rebuild that adds th | `SURREAL_ROCKSDB_MAX_WRITE_BUFFER_NUMBER` | `32` | Maximum number of write buffers in memory at once. | | `SURREAL_ROCKSDB_BACKGROUND_THREADS` | `4` | Number of background compaction threads. | -Give the Nexus PDP container a memory limit of 4 GiB for a few million facts and more for a larger data set. Never set the limit below 2.5 GiB: below 2304 MiB, the embedded database can stop applying changes without reporting it. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint). +Give the Nexus PDP container a memory limit of 4 GiB to start, and more for a large data set. Never set the limit below 2.5 GiB: below 2304 MiB, the embedded database can stop applying changes without reporting it. See [Nexus PDP resource footprint](/concepts/pdp/nexus-pdp-how-it-works#resource-footprint). ## Next steps \{#related-documentation} diff --git a/docs/concepts/pdp/nexus-pdp-deployment.mdx b/docs/concepts/pdp/nexus-pdp-deployment.mdx index d4422500..f9aa47fe 100644 --- a/docs/concepts/pdp/nexus-pdp-deployment.mdx +++ b/docs/concepts/pdp/nexus-pdp-deployment.mdx @@ -22,13 +22,13 @@ Nexus PDP has operational requirements that the container PDP (the Edge PDP imag | Image | `permitio/nexus-pdp`, pinned by digest (`permitio/nexus-pdp@sha256:`) | `permitio/nexus-pdp` has no `latest` tag, so a pull without a tag fails. Tags can move: `0-beta` moves to each new beta or release build, and a numbered beta tag name can be published again for a later build. A new build can rename configuration variables. See [Nexus PDP configuration reference](/concepts/pdp/nexus-pdp-configuration). | | One container per environment | Set `PDP_API_KEY` to the Nexus PDP API key of one Permit environment | Nexus PDP has no multi-environment mode. The Nexus PDP API key binds the container to exactly one environment. | | Persistent storage | Mount a persistent volume at `/var/lib/edge-pdp`, which holds the embedded database and the event store at default paths | On ephemeral storage, every restart runs a full cold start with a snapshot transfer of your whole data set. | -| Memory limit (`resources.limits.memory`) | 4 GiB to start, and never less than 2.5 GiB; about 7 GiB at 10 million [facts](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint) and 12 GiB at 20 million (estimates that allow for a rebuild) | Below 2304 MiB (2.25 GiB), the embedded database can stop applying changes under a sustained write load without reporting it; raise the memory limit and restart the container to recover. A large data set needs more memory during a cold start than afterward, and the container is killed if the cold start passes the limit. The embedded database sizes its cache from the memory limit; without a limit, it sizes the cache from the node's memory. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint). | +| Memory limit (`resources.limits.memory`) | 4 GiB to start, more for a large data set, and never less than 2.5 GiB | Below 2304 MiB (2.25 GiB), the embedded database can stop applying changes under a sustained write load without reporting it; raise the memory limit and restart the container to recover. A large data set needs more memory during a cold start or a rebuild than afterward, and the container is killed if either passes the limit. The embedded database sizes its cache from the memory limit; without a limit, it sizes the cache from the node's memory. See [Nexus PDP resource footprint](/concepts/pdp/nexus-pdp-how-it-works#resource-footprint). | | Volume permissions | The volume is writable by user ID and group ID `10001` | Nexus PDP runs as the non-root user `10001` and cannot write its database or event store. | | Exposed ports | `7000` (authorization API) and `7001` (health) only | Other ports bind to loopback inside the container and are not reachable from outside. | | Termination grace period | `terminationGracePeriodSeconds: 40` or higher | 40 seconds is the sum of the default shutdown budgets `EDGE_DRAIN_TIMEOUT_SECS` (10) and `EDGE_CHILD_TERMINATION_TIMEOUT_SECS` (30). With a shorter grace period, Kubernetes kills the NATS leaf node before its on-disk event store finishes flushing. | | Liveness probe | `GET /health` on port `7001` | A liveness probe on port `7000` fails during the cold start and restarts the container. See [Health and readiness](#health-and-readiness). | | Readiness probe | `GET /health/ready` on port `7001` | Traffic reaches a Nexus PDP that cannot answer yet. | -| Startup probe | `GET /health/ready` on port `7001`, with `failureThreshold` × `periodSeconds` longer than your cold start plus headroom (measured on 8 CPUs: 68 s at 10.8 million facts, 151 s at 20 million). Fewer CPUs take longer, so time a cold start of your own data set before you set the probe | A cold start takes longer as your data grows. A short startup probe restarts a first boot that is still loading data. | +| Startup probe | `GET /health/ready` on port `7001`, with `failureThreshold` × `periodSeconds` longer than your cold start plus headroom. A cold start takes longer with more data and fewer CPUs, so time one with your own data set before you set the probe | A cold start takes longer as your data grows. A short startup probe restarts a first boot that is still loading data. | ## Run Nexus PDP with Docker diff --git a/docs/concepts/pdp/nexus-pdp-how-it-works.mdx b/docs/concepts/pdp/nexus-pdp-how-it-works.mdx index 302c5b32..965b4253 100644 --- a/docs/concepts/pdp/nexus-pdp-how-it-works.mdx +++ b/docs/concepts/pdp/nexus-pdp-how-it-works.mdx @@ -120,46 +120,24 @@ Nexus PDP stores authorization data in a different place than the container PDP, | | Container PDP (`pdp-v2`) | Nexus PDP (`nexus-pdp`) | | --- | --- | --- | | Authorization data | In OPA's in-memory document | On disk, in an embedded database | -| Memory as data grows | Grows with your data set | Grows with your data set, and peaks during the cold start of a large data set | +| Memory as data grows | Grows with your data set | Grows with your data set, and peaks while it loads a snapshot | | Processes | Rust API server, Python OPAL client (Horizon), and OPA | Rust binary, NATS leaf node, and OPA | | Python runtime | Required, for the Open Policy Administration Layer (OPAL) client | Not present | | Persistent storage | Not required | Required | -On the container PDP, your authorization data is raw JSON in memory. The [OPA policy performance documentation](https://www.openpolicyagent.org/docs/policy-performance) states that raw JSON in OPA uses about 20 times the memory of the same data in a compact on-disk form. On Nexus PDP, the data set lives on disk in a compact form. Memory still grows with the data set, and on a large data set it peaks during the cold start, while Nexus PDP loads the snapshot: see [Measured memory and cold start](#measured-footprint). +On the container PDP, your authorization data is raw JSON in memory. The [OPA policy performance documentation](https://www.openpolicyagent.org/docs/policy-performance) states that raw JSON in OPA uses about 20 times the memory of the same data in a compact on-disk form. On Nexus PDP, the data set lives on disk in a compact form. Memory still grows with the data set, and peaks while Nexus PDP loads a snapshot: during a cold start, and during a [rebuild](#self-healing), when the old database keeps serving with its cache while the new one loads. :::caution Never give Nexus PDP less than 2.5 GiB of memory Below 2304 MiB (2.25 GiB), the embedded database can stop applying changes for good under a sustained write load, such as applying the changes that arrived during a cold start, or a burst of changes. Nexus PDP then keeps answering from the data it already has, so later changes to your policy data don't take effect. As of `0.7.0-beta.26`, Nexus PDP doesn't report this: `/health` stays up and nothing is logged. With 2 CPUs or fewer, the whole process can stop answering instead, and a liveness probe on `/health` restarts it. To recover, restart the container and raise its memory limit. A load test of permission checks alone doesn't show the problem, because checks don't write. 2304 MiB is the point below which writes can stall, not a size to run at. Memory limits in decimal units count as written: `2400M` (2.4 billion bytes) is below it. -Give the Nexus PDP container a memory limit of 4 GiB to start, and more for a large data set: see [Measured memory and cold start](#measured-footprint). The embedded database sizes its cache and write buffers from the container's memory limit; without a limit, it sizes them from the node's memory. See [Nexus PDP storage engine settings](/concepts/pdp/nexus-pdp-configuration#storage-engine). Before you choose a size, load test against your own data set, and include a cold start and a sustained burst of updates in the test. +Give the Nexus PDP container a memory limit of 4 GiB to start, and more for a large data set. The embedded database sizes its cache and write buffers from the container's memory limit; without a limit, it sizes them from the node's memory. See [Nexus PDP storage engine settings](/concepts/pdp/nexus-pdp-configuration#storage-engine). Before you choose a size, load test against your own data set, and include a cold start, a rebuild, and a sustained burst of updates in the test. ::: -### Measured memory and cold start (early access) \{#measured-footprint} - -Permit measured these figures in October 2026 on Nexus PDP `0.7.0-beta.26` (`permitio/nexus-pdp@sha256:fa7012adc2f2981829f914f211f599bc3e127625ec764adcb0056dbe54ff1bdb`), an early-access build, with default storage-engine settings. The data set was one Google Drive-style ReBAC model (tenants, groups, folders, and documents, nested up to five levels), loaded in three steps. Facts are the rows Permit stores for an environment: users, tenants, resource instances, role assignments, and relationship tuples. Treat the figures as a starting point, not a guarantee: later builds can change them, and your policy and data shape change them too. - -| Facts | Users / resource instances | CPUs | Cold start, to correct data | Peak memory during the cold start (at least) | Memory while serving checks | Snapshot size | -| --- | --- | --- | --- | --- | --- | --- | -| 2.0M | 182K / 402K | 2 | 24 s | Not recorded | Up to 1.5 GiB | 0.26 GiB | -| 10.8M | 1.0M / 2.2M | 8 | 68 s | ~3.9 GiB | 2.1–2.9 GiB | 1.35 GiB | -| 20.0M | 1.85M / 4.1M | 8 | 151 s | ~6.5 GiB | 1.7–3.2 GiB | 2.67 GiB | - -Peak memory is the highest container working set sampled every 10 seconds, so the real peak can be higher. The runs had memory limits of 8 GiB (2.0M) and 32 GiB (10.8M and 20.0M). Because the embedded database sizes its cache from the memory limit, memory while serving can differ under a smaller limit. The limits recommended below are estimates that allow for a rebuild, and were not tested as limits. - -- **Memory limit by data-set size.** Never set the memory limit below 2.5 GiB, whatever the size of your data set: below 2304 MiB (2.25 GiB), the embedded database can stop applying changes without reporting it. 4 GiB covers a few million facts. For a larger data set, give at least four times the snapshot size plus 1 GiB, rounded up: about 7 GiB at 10 million facts and 12 GiB at 20 million. That leaves room for a [rebuild](#self-healing), which loads the snapshot while the old database keeps serving with a cache of up to half the memory limit minus 1 GiB. These limits are estimates: Permit has not measured a rebuild or larger data sets, so load test before you choose a size. -- **On a large data set, memory peaks during a cold start.** Loading the snapshot takes about twice the snapshot's size on top of the rest of Nexus PDP. At 10.8 million and 20 million facts, the peak was about twice the snapshot plus 1.2 GiB. Set the memory limit above that peak with headroom, or the container is killed while it loads. A [rebuild](#self-healing) runs the same load while Nexus PDP keeps serving from the old database, so it needs more memory than a first boot; the limits in the previous bullet allow for it. Permit has not measured a rebuild. -- **Cold start time grows with the data set.** On 8 CPUs, it took 68 s at 10.8 million facts and 151 s at 20 million: 1.85 times the facts took 2.2 times as long. Set the startup probe from the cold start time for your size, with headroom. On a new Kubernetes node, the pod also waits about a minute for the node and the image pull before the container starts. That wait counts against rollout timeouts, not the startup probe. -- **The snapshot grows linearly**, at about 0.13 GiB for each million facts. -- **Disk.** After the cold start, the volume used 0.33 GiB at 2.0 million facts and 1.6 GiB at 10.8 million. At 20 million facts, it used up to 5.8 GiB during the cold start and 2.8 GiB after it. No rebuild was measured. - -For comparison, the container PDP's [memory estimate](/how-to/deploy/deploy-to-production#scaling-memory) of 6 KB per user, resource instance, and tenant gives about 36 GB for the 20 million facts data set. That is an estimate, not a measurement of the container PDP on the same data. - -These figures cover memory, disk, and cold start only. They are not latency or throughput figures: per-check latency depends on how many resources each user reaches in your policy and data, so measure it against your own data set. The [caution on throughput](#running-at-high-volume) still applies. - ### Disk -Two Nexus PDP paths must be on persistent storage: the embedded database (`EDGE_DB_PATH`) and the event store (`EDGE_DATA_DIR`). At the default paths, both are on the volume mounted at `/var/lib/edge-pdp`. Size the volume from your snapshot size, plus headroom. In the [measured data set](#measured-footprint), the snapshot was about 0.13 GiB for each million facts. +Two Nexus PDP paths must be on persistent storage: the embedded database (`EDGE_DB_PATH`) and the event store (`EDGE_DATA_DIR`). At the default paths, both are on the volume mounted at `/var/lib/edge-pdp`. Size the volume from your snapshot size, plus headroom. - **During a cold start**, the volume holds the downloaded snapshot, an unpacked copy of its files, and the database built from them, all at once: about three times the snapshot. If a cold start is killed before it finishes, for example by the memory limit or a short startup probe, its copies can stay on the volume, and each new attempt writes new ones. If restarts fill the volume, fix the cause, then empty the volume so the next start runs a clean cold start. - **During a [rebuild](#self-healing)**, the old database also stays on the volume until Nexus PDP swaps in the new copy: about four times the snapshot. Give the volume at least four times the snapshot size. diff --git a/docs/concepts/pdp/nexus-pdp.mdx b/docs/concepts/pdp/nexus-pdp.mdx index 1cc7cbe2..bd5de6e6 100644 --- a/docs/concepts/pdp/nexus-pdp.mdx +++ b/docs/concepts/pdp/nexus-pdp.mdx @@ -47,7 +47,7 @@ Nexus PDP stores relationship data in an embedded database built for graph trave The container PDP holds all authorization data in memory, so a large environment needs a large container PDP. The [OPA policy performance documentation](https://www.openpolicyagent.org/docs/policy-performance) states that raw JSON loaded into OPA uses about 20 times the memory of the same data in a compact serialized form. -Nexus PDP stores the compact form on disk and keeps a cache in memory, sized from the container's memory limit. Its memory still grows with your data set, and peaks during the cold start of a large data set. See [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint) for figures by data-set size. +Nexus PDP stores the compact form on disk and keeps a cache in memory, sized from the container's memory limit. Its memory still grows with your data set, and peaks while it loads a snapshot. See [Nexus PDP resource footprint](/concepts/pdp/nexus-pdp-how-it-works#resource-footprint). ### Container PDP updates need a second request \{#updates-cost-a-round-trip} @@ -58,7 +58,7 @@ Nexus PDP receives each update with the changed data inside the message. The con ## What Nexus PDP provides \{#what-it-gives-you} - **Decisions inside the container.** Every hop in a decision runs over loopback or reads local disk. No authorization query depends on reaching Permit. -- **Less memory for the same data.** The data set lives on disk, and a cache sized from the container's memory limit holds the hot part of it in memory. +- **Data on disk.** The data set lives on disk, and a cache sized from the container's memory limit holds the hot part of it in memory. - **Sync that resumes after a disconnection.** The control plane keeps changes for each Nexus PDP until that PDP acknowledges them, so a reconnecting Nexus PDP does not fetch its data set again. - **Decisions during control-plane outages.** When the control plane is unreachable, Nexus PDP keeps answering from its local copy and reports how stale that copy is on its health endpoint. diff --git a/docs/concepts/pdp/overview.mdx b/docs/concepts/pdp/overview.mdx index 564270bd..5f9a6f6b 100644 --- a/docs/concepts/pdp/overview.mdx +++ b/docs/concepts/pdp/overview.mdx @@ -50,7 +50,7 @@ The Edge PDP bundles three components in one container: Open Policy Agent (OPA), Nexus PDP differs from the Edge PDP in four ways: -- **Data on disk, not in memory.** The Edge PDP holds your data in OPA's in-memory JSON document, so a large environment needs a large PDP. Nexus PDP stores the data in an embedded database on disk and keeps a cache in memory, sized from the container's memory limit. Nexus PDP memory still grows with your data set, and peaks during the cold start of a large data set: see [Measured memory and cold start](/concepts/pdp/nexus-pdp-how-it-works#measured-footprint). +- **Data on disk, not in memory.** The Edge PDP holds your data in OPA's in-memory JSON document, so a large environment needs a large PDP. Nexus PDP stores the data in an embedded database on disk and keeps a cache in memory, sized from the container's memory limit. Nexus PDP memory still grows with your data set, and peaks while it loads a snapshot: see [Nexus PDP resource footprint](/concepts/pdp/nexus-pdp-how-it-works#resource-footprint). - **A store built for relationship queries.** Nexus PDP keeps relationship data in a database built for graph traversal, which OPA queries over loopback while it evaluates policy. - **Sync that resumes after disconnection.** Each update contains the changed data, and the control plane keeps each Nexus PDP's unacknowledged changes until that PDP applies them. A reconnecting Nexus PDP does not fetch its data again. - **Decisions during control-plane outages.** If the control plane is unreachable, Nexus PDP keeps answering from its local copy and reports how stale the copy is.