Skip to content
18 changes: 14 additions & 4 deletions docs/concepts/pdp/nexus-pdp-configuration.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -71,16 +71,26 @@ Set the Kubernetes `terminationGracePeriodSeconds` to at least `EDGE_DRAIN_TIMEO

## Storage engine

The Nexus PDP embedded database (SurrealDB on RocksDB) uses storage-engine defaults sized for large workloads, not for a small container. These variables have the largest effect on memory use:
The Nexus PDP embedded database is SurrealDB on RocksDB. As of `0.7.0-beta.26`, the database that serves checks sizes its memory from the container's memory limit, not from environment variables:

| Variable | Default | Effect |
| Setting | Size |
| --- | --- |
| Block cache (the read cache) | Half the memory limit minus 1 GiB, and at least 16 MiB |
| Size of each write buffer | 32 MiB below a 1 GiB limit, 64 MiB below 16 GiB, and 128 MiB from 16 GiB |
| Write buffers in memory at once | 2 below a 4 GiB limit, 4 below 16 GiB, 8 below 64 GiB, and 32 from 64 GiB |

If the container has no memory limit, the database sizes these from the node's total memory instead. Set a memory limit, for example `resources.limits.memory` on Kubernetes.

These variables apply only to the step of a cold start or a rebuild that adds the snapshot files to a new database, and they have little effect on memory even there: that step adds whole files instead of writing through the write buffers, and its memory peak comes from holding the snapshot. They don't change the memory of the database that serves checks:

| Variable | Default | Setting it controls |
| --- | --- | --- |
| `SURREAL_ROCKSDB_BLOCK_CACHE_SIZE` | `536870912` (512 MiB) | Size of the read cache. The largest contributor to resident memory. |
| `SURREAL_ROCKSDB_BLOCK_CACHE_SIZE` | `536870912` (512 MiB) | Size of the read cache. |
| `SURREAL_ROCKSDB_WRITE_BUFFER_SIZE` | `268435456` (256 MiB) | Size of each in-memory write buffer before RocksDB flushes it to disk. |
| `SURREAL_ROCKSDB_MAX_WRITE_BUFFER_NUMBER` | `32` | Maximum number of write buffers in memory at once. |
| `SURREAL_ROCKSDB_BACKGROUND_THREADS` | `4` | Number of background compaction threads. |

At these defaults, a container with a memory limit of a few hundred MiB is killed for running out of memory (OOM) at startup. Start with 4 GiB of memory, then lower these values while you load test against your own data set. See [Nexus PDP resource footprint](/concepts/pdp/nexus-pdp-how-it-works#resource-footprint).
Give the Nexus PDP container a memory limit of 4 GiB to start, and more for a large data set. Never set the limit below 2.5 GiB: below 2304 MiB, the embedded database can stop applying changes without reporting it. See [Nexus PDP resource footprint](/concepts/pdp/nexus-pdp-how-it-works#resource-footprint).

## Next steps \{#related-documentation}

Expand Down
14 changes: 10 additions & 4 deletions docs/concepts/pdp/nexus-pdp-deployment.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -22,13 +22,13 @@ Nexus PDP has operational requirements that the container PDP (the Edge PDP imag
| Image | `permitio/nexus-pdp`, pinned by digest (`permitio/nexus-pdp@sha256:<digest>`) | `permitio/nexus-pdp` has no `latest` tag, so a pull without a tag fails. Tags can move: `0-beta` moves to each new beta or release build, and a numbered beta tag name can be published again for a later build. A new build can rename configuration variables. See [Nexus PDP configuration reference](/concepts/pdp/nexus-pdp-configuration). |
| One container per environment | Set `PDP_API_KEY` to the Nexus PDP API key of one Permit environment | Nexus PDP has no multi-environment mode. The Nexus PDP API key binds the container to exactly one environment. |
| Persistent storage | Mount a persistent volume at `/var/lib/edge-pdp`, which holds the embedded database and the event store at default paths | On ephemeral storage, every restart runs a full cold start with a snapshot transfer of your whole data set. |
| Memory | 4 GiB to start | At default storage-engine settings, a container with a few hundred MiB is killed for running out of memory (OOM) at startup. See [Nexus PDP resource footprint](/concepts/pdp/nexus-pdp-how-it-works#resource-footprint). |
| Memory limit (`resources.limits.memory`) | 4 GiB to start, more for a large data set, and never less than 2.5 GiB | Below 2304 MiB (2.25 GiB), the embedded database can stop applying changes under a sustained write load without reporting it; raise the memory limit and restart the container to recover. A large data set needs more memory during a cold start or a rebuild than afterward, and the container is killed if either passes the limit. The embedded database sizes its cache from the memory limit; without a limit, it sizes the cache from the node's memory. See [Nexus PDP resource footprint](/concepts/pdp/nexus-pdp-how-it-works#resource-footprint). |
| Volume permissions | The volume is writable by user ID and group ID `10001` | Nexus PDP runs as the non-root user `10001` and cannot write its database or event store. |
| Exposed ports | `7000` (authorization API) and `7001` (health) only | Other ports bind to loopback inside the container and are not reachable from outside. |
| Termination grace period | `terminationGracePeriodSeconds: 40` or higher | 40 seconds is the sum of the default shutdown budgets `EDGE_DRAIN_TIMEOUT_SECS` (10) and `EDGE_CHILD_TERMINATION_TIMEOUT_SECS` (30). With a shorter grace period, Kubernetes kills the NATS leaf node before its on-disk event store finishes flushing. |
| Liveness probe | `GET /health` on port `7001` | A liveness probe on port `7000` fails during the cold start and restarts the container. See [Health and readiness](#health-and-readiness). |
| Readiness probe | `GET /health/ready` on port `7001` | Traffic reaches a Nexus PDP that cannot answer yet. |
| Startup probe | Enough time for a cold start of your data set | A cold start takes longer as your data grows. A short startup probe restarts a first boot that is still loading data. |
| Startup probe | `GET /health/ready` on port `7001`, with `failureThreshold` × `periodSeconds` longer than your cold start plus headroom. A cold start takes longer with more data and fewer CPUs, so time one with your own data set before you set the probe | A cold start takes longer as your data grows. A short startup probe restarts a first boot that is still loading data. |

## Run Nexus PDP with Docker

Expand All @@ -45,14 +45,14 @@ docker run -d --name nexus-pdp \

## Kubernetes pod settings for Nexus PDP

This pod-spec excerpt sets the ports, probes, memory request, volume, and shutdown budget from the requirements table. `fsGroup: 10001` makes the mounted volume writable by the non-root user that Nexus PDP runs as. Create the `nexus-pdp-api-key` secret with the environment's Nexus PDP API key under the key `PDP_API_KEY`, which the pod spec reads, reusing the shell variable from the Docker step:
This pod-spec excerpt sets the ports, probes, memory request and limit, volume, and shutdown budget from the requirements table. `fsGroup: 10001` makes the mounted volume writable by the non-root user that Nexus PDP runs as. Create the `nexus-pdp-api-key` secret with the environment's Nexus PDP API key under the key `PDP_API_KEY`, which the pod spec reads, reusing the shell variable from the Docker step:

```bash
kubectl create secret generic nexus-pdp-api-key \
--from-literal=PDP_API_KEY="$NEXUS_PDP_API_KEY"
```

Then size the `nexus-pdp-data` claim for your data set, and add a startup probe with enough time for a cold start:
Then size the `nexus-pdp-data` claim for your data set (see [Disk](/concepts/pdp/nexus-pdp-how-it-works#disk)), and add a startup probe with enough time for a cold start:

```yaml
terminationGracePeriodSeconds: 40
Expand All @@ -67,6 +67,8 @@ containers:
resources:
requests:
memory: 4Gi
limits:
memory: 4Gi
env:
- name: PDP_API_KEY
valueFrom:
Expand All @@ -77,6 +79,10 @@ containers:
httpGet: { path: /health, port: 7001 }
readinessProbe:
httpGet: { path: /health/ready, port: 7001 }
startupProbe:
httpGet: { path: /health/ready, port: 7001 }
periodSeconds: 10
failureThreshold: 30 # 300 s: set from a timed cold start of your data set
volumeMounts:
- name: nexus-data
mountPath: /var/lib/edge-pdp
Expand Down
20 changes: 13 additions & 7 deletions docs/concepts/pdp/nexus-pdp-how-it-works.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ The NATS leaf node inside the Nexus PDP container keeps a persistent connection

| Plane | Carries |
| --- | --- |
| **Change stream** | Transactions of authorization data: users, tenants, resource instances, and relationship tuples |
| **Change stream** | Transactions of authorization data: users, tenants, resource instances, role assignments, and relationship tuples |
| **Policy files** | The compiled Rego bundle that Open Policy Agent (OPA) evaluates |
| **Policy schema** | Role and permission definitions from your policy |
| **Snapshot** | A one-time bulk transfer of the whole data set, used only during a cold start |
Expand Down Expand Up @@ -120,22 +120,28 @@ Nexus PDP stores authorization data in a different place than the container PDP,
| | Container PDP (`pdp-v2`) | Nexus PDP (`nexus-pdp`) |
| --- | --- | --- |
| Authorization data | In OPA's in-memory document | On disk, in an embedded database |
| Memory as data grows | Grows with your data set | Limited by a cache size you configure |
| Memory as data grows | Grows with your data set | Grows with your data set, and peaks while it loads a snapshot |
| Processes | Rust API server, Python OPAL client (Horizon), and OPA | Rust binary, NATS leaf node, and OPA |
| Python runtime | Required, for the Open Policy Administration Layer (OPAL) client | Not present |
| Persistent storage | Not required | Required |

On the container PDP, your authorization data is raw JSON in memory. The [OPA policy performance documentation](https://www.openpolicyagent.org/docs/policy-performance) states that raw JSON in OPA uses about 20 times the memory of the same data in a compact on-disk form. On Nexus PDP, the data set lives on disk in a compact form, and memory depends on the storage engine's cache sizes, which you set.
On the container PDP, your authorization data is raw JSON in memory. The [OPA policy performance documentation](https://www.openpolicyagent.org/docs/policy-performance) states that raw JSON in OPA uses about 20 times the memory of the same data in a compact on-disk form. On Nexus PDP, the data set lives on disk in a compact form. Memory still grows with the data set, and peaks while Nexus PDP loads a snapshot: during a cold start, and during a [rebuild](#self-healing), when the old database keeps serving with its cache while the new one loads.

:::caution Start Nexus PDP with 4 GiB of memory
The default storage-engine settings of Nexus PDP target large workloads: a 512 MiB block cache and 256 MiB write buffers. If you give the container a memory limit of a few hundred MiB, the container is killed for running out of memory (OOM) at startup. Give the Nexus PDP container 4 GiB of memory to start.
:::caution Never give Nexus PDP less than 2.5 GiB of memory
Below 2304 MiB (2.25 GiB), the embedded database can stop applying changes for good under a sustained write load, such as applying the changes that arrived during a cold start, or a burst of changes. Nexus PDP then keeps answering from the data it already has, so later changes to your policy data don't take effect. As of `0.7.0-beta.26`, Nexus PDP doesn't report this: `/health` stays up and nothing is logged. With 2 CPUs or fewer, the whole process can stop answering instead, and a liveness probe on `/health` restarts it. To recover, restart the container and raise its memory limit. A load test of permission checks alone doesn't show the problem, because checks don't write.

To run Nexus PDP with less memory, lower the cache and write-buffer sizes in [Nexus PDP storage engine settings](/concepts/pdp/nexus-pdp-configuration#storage-engine), then load test against your own data set before you choose a size.
2304 MiB is the point below which writes can stall, not a size to run at. Memory limits in decimal units count as written: `2400M` (2.4 billion bytes) is below it.

Give the Nexus PDP container a memory limit of 4 GiB to start, and more for a large data set. The embedded database sizes its cache and write buffers from the container's memory limit; without a limit, it sizes them from the node's memory. See [Nexus PDP storage engine settings](/concepts/pdp/nexus-pdp-configuration#storage-engine). Before you choose a size, load test against your own data set, and include a cold start, a rebuild, and a sustained burst of updates in the test.
:::

### Disk

Two Nexus PDP paths must be on persistent storage: the embedded database (`EDGE_DB_PATH`) and the event store (`EDGE_DATA_DIR`). Plan for about twice the size of your data set, plus headroom. During a rebuild, the old and new database copies exist on disk together until Nexus PDP swaps in the new copy.
Two Nexus PDP paths must be on persistent storage: the embedded database (`EDGE_DB_PATH`) and the event store (`EDGE_DATA_DIR`). At the default paths, both are on the volume mounted at `/var/lib/edge-pdp`. Size the volume from your snapshot size, plus headroom.

- **During a cold start**, the volume holds the downloaded snapshot, an unpacked copy of its files, and the database built from them, all at once: about three times the snapshot. If a cold start is killed before it finishes, for example by the memory limit or a short startup probe, its copies can stay on the volume, and each new attempt writes new ones. If restarts fill the volume, fix the cause, then empty the volume so the next start runs a clean cold start.
- **During a [rebuild](#self-healing)**, the old database also stays on the volume until Nexus PDP swaps in the new copy: about four times the snapshot. Give the volume at least four times the snapshot size.
- **The event store** keeps the changes Nexus PDP receives, with no age or size limit. It grows with your update volume over time, not with the data set, so monitor free space on the volume.

## Next steps \{#related-documentation}

Expand Down
4 changes: 2 additions & 2 deletions docs/concepts/pdp/nexus-pdp.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,7 @@ Nexus PDP stores relationship data in an embedded database built for graph trave

The container PDP holds all authorization data in memory, so a large environment needs a large container PDP. The [OPA policy performance documentation](https://www.openpolicyagent.org/docs/policy-performance) states that raw JSON loaded into OPA uses about 20 times the memory of the same data in a compact serialized form.

Nexus PDP stores the compact form on disk and keeps a cache of configurable size in memory. You set the memory budget instead of deriving it from the size of your data.
Nexus PDP stores the compact form on disk and keeps a cache in memory, sized from the container's memory limit. Its memory still grows with your data set, and peaks while it loads a snapshot. See [Nexus PDP resource footprint](/concepts/pdp/nexus-pdp-how-it-works#resource-footprint).

### Container PDP updates need a second request \{#updates-cost-a-round-trip}

Expand All @@ -58,7 +58,7 @@ Nexus PDP receives each update with the changed data inside the message. The con
## What Nexus PDP provides \{#what-it-gives-you}

- **Decisions inside the container.** Every hop in a decision runs over loopback or reads local disk. No authorization query depends on reaching Permit.
- **Memory you configure.** The data set lives on disk, and a configurable cache limits resident memory.
- **Data on disk.** The data set lives on disk, and a cache sized from the container's memory limit holds the hot part of it in memory.
- **Sync that resumes after a disconnection.** The control plane keeps changes for each Nexus PDP until that PDP acknowledges them, so a reconnecting Nexus PDP does not fetch its data set again.
- **Decisions during control-plane outages.** When the control plane is unreachable, Nexus PDP keeps answering from its local copy and reports how stale that copy is on its health endpoint.

Expand Down
2 changes: 1 addition & 1 deletion docs/concepts/pdp/overview.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@ The Edge PDP bundles three components in one container: Open Policy Agent (OPA),

Nexus PDP differs from the Edge PDP in four ways:

- **Data on disk, not in memory.** The Edge PDP holds your data in OPA's in-memory JSON document, so a large environment needs a large PDP. Nexus PDP stores the data in an embedded database on disk and keeps a cache of configurable size in memory.
- **Data on disk, not in memory.** The Edge PDP holds your data in OPA's in-memory JSON document, so a large environment needs a large PDP. Nexus PDP stores the data in an embedded database on disk and keeps a cache in memory, sized from the container's memory limit. Nexus PDP memory still grows with your data set, and peaks while it loads a snapshot: see [Nexus PDP resource footprint](/concepts/pdp/nexus-pdp-how-it-works#resource-footprint).
- **A store built for relationship queries.** Nexus PDP keeps relationship data in a database built for graph traversal, which OPA queries over loopback while it evaluates policy.
- **Sync that resumes after disconnection.** Each update contains the changed data, and the control plane keeps each Nexus PDP's unacknowledged changes until that PDP applies them. A reconnecting Nexus PDP does not fetch its data again.
- **Decisions during control-plane outages.** If the control plane is unreachable, Nexus PDP keeps answering from its local copy and reports how stale the copy is.
Expand Down
Loading