From 30887a396e633e1b5c93d528682b9ed4e778b616 Mon Sep 17 00:00:00 2001 From: Viktor Voytovych Date: Tue, 29 Sep 2026 14:08:09 +0300 Subject: [PATCH 1/5] docs: Propose one writer per zone for DNS records Two record sets can claim one owner name. The agent writes one record set at a time, so it cannot choose between them, and when either one lets the name go, the agent deletes the name the other still claims (#188). The enhancement proposes one reconcile per zone, which builds the zone from every record set that references it and writes the difference. It is rolled out in small releases, each verified locally and on staging. Related to #188 --- .../dnsrecordset-per-zone-writer.md | 178 ++++++++++++++++++ 1 file changed, 178 insertions(+) create mode 100644 docs/enhancements/dnsrecordset-per-zone-writer.md diff --git a/docs/enhancements/dnsrecordset-per-zone-writer.md b/docs/enhancements/dnsrecordset-per-zone-writer.md new file mode 100644 index 0000000..464f73b --- /dev/null +++ b/docs/enhancements/dnsrecordset-per-zone-writer.md @@ -0,0 +1,178 @@ +# Enhancement: One Writer per Zone for DNS Records + +## Summary + +Two `DNSRecordSet` objects can claim the same owner name in one zone. The agent writes records one record +set at a time, so it never sees both claims together, and it cannot decide between them. When one of the +two lets the name go, the agent deletes the name, although the other still claims it +([#188](https://github.com/datum-cloud/dns-operator/issues/188)). We propose that the agent writes one +zone at a time instead: it reads every record set of the zone, picks one holder for each name, and writes +the difference in one pass. + +## Terms + +- A record set **claims** an owner name when it lists that name. +- Of the record sets that claim a name with one record type, the **holder** is the one whose records are + served: the oldest, and at one creation time, the lower name. +- A record set **releases** a name when it is deleted, or when the name leaves its spec. +- The agent stores an **ownership note** with each record it writes, naming the record set that wrote it. + +## Motivation + +[`record-ownership.md`](https://github.com/datum-cloud/dns-operator/blob/main/docs/architecture/record-ownership.md) states the rule. When a record set releases +a name, *"the remaining claimants are re-elected … If no claimant remains, the backend record set for +that name is deleted."* Until `v0.7.0` an agent reconciler handled one name at a time, applied this rule, +and told each record set that lost the election `NotOwner`. + +[#82](https://github.com/datum-cloud/dns-operator/pull/82) replaced it with one reconcile per record set +([design](https://github.com/datum-cloud/dns-operator/blob/main/docs/enhancements/dnsrecordset-reconcile-consolidation.md)). That fixed a delete that did not survive a restart +([#58](https://github.com/datum-cloud/dns-operator/issues/58)) and a queue that fanned every zone event out +to every name ([#59](https://github.com/datum-cloud/dns-operator/issues/59)). But a reconcile that sees +one record set cannot choose among several. So three behaviours went with the old reconciler: only the +holder's records are written, the others report `NotOwner`, and a released name passes to the next +claimant. [#86](https://github.com/datum-cloud/dns-operator/issues/86) proposes removing the old +reconciler, which nothing runs. + +The admission webhook refuses a second claim on a name. It cannot make duplicates impossible: + +- it fails open on purpose: in the outage of [#69](https://github.com/datum-cloud/dns-operator/issues/69), + a webhook set to fail closed refused every record set write while no running release could serve it; +- two creates at the same moment both pass it; +- it does not see claims that existed before it. + +[#83](https://github.com/datum-cloud/dns-operator/issues/83) found such pairs in production. So the agent +still has to decide what gets through, as a Gateway API controller does for conflicting routes. + +## Goals + +- For every name in a zone, the holder's records are served, the other claimants report `NotOwner`, a + released name passes to the next claimant, and it is deleted when no claimant is left. +- Each step ships as its own release, and the release before it is the rollback. Each step changes at + most one thing that users see, so a problem points at one step. +- In a zone without two claims on one name, the records served do not change, with one exception: + records that carry an ownership note although no record set claims them any more. These are leftovers + of failures such as [#58](https://github.com/datum-cloud/dns-operator/issues/58), and the first pass removes them, within a limit. + +## Non-Goals + +- Removing records that no record set declares and that carry no ownership note. This is DNSControl's + default, which its [`NO_PURGE`](https://github.com/StackExchange/dnscontrol/blob/main/documentation/language-reference/domain-modifiers/NO_PURGE.md) + switch turns off for zones that other systems also write. We would need a per-zone switch of our own + first. +- Replacing the ownership notes. Until a zone is fully declared, they are the only mark of which records + the agent may delete. +- The SOA serial, and the zone controller's writes of the SOA and NS records. Moving those into this + writer is a separate decision. + +## Background + +- **A Kubernetes work queue never processes one item in two workers at once** + ([client-go](https://pkg.go.dev/k8s.io/client-go/util/workqueue): *"a single item will not be processed + multiple times concurrently"*). With the zone as the item, each zone has one writer, and no lock is + needed. +- **Gateway API breaks conflicts by the oldest creation timestamp, then by namespace and name** + ([`gateway_types.go`](https://github.com/kubernetes-sigs/gateway-api/blob/main/apis/v1/gateway_types.go)). + Our election uses the same order. +- **DNSControl and octoDNS both build a zone from its declared state and apply the difference.** This + proposal does the same, limited to the records the agent owns. + +## Options + +1. **Hand a released name over inside today's per-record-set reconcile.** We built and reviewed this. It + needs a lock per zone, a separate release path, and the claimants passed down into the PowerDNS client, + and it still reports no `NotOwner`. It adds code to work around the unit of work. +2. **Rely on the webhook alone.** It cannot close the gaps listed above. +3. **One reconcile per zone and record type.** The item is smaller, but some rules cross types, such as a + CNAME that may not share a name with other records, and those need the whole zone. +4. **One reconcile per zone.** Chosen. The rule is about several record sets, so the unit of work holds + them all. A prototype built on 2026-09-29 passed every unit, integration and end-to-end suite, and + added about as many lines of code as it removed (+288 / −295). On that date the old reconciler + ([#86](https://github.com/datum-cloud/dns-operator/issues/86)) held about 1,350 more lines, which + step 5 removes. + +## Design + +**What is queued.** A zone is queued when one of its record sets is created, deleted, changes its spec, +or starts deleting, and when the zone itself becomes programmed. When a record set moves to another zone, +both zones are queued, so the old zone releases its names. Writes to status and metadata are not queued: +a reconcile that queued its own status writes would run in a loop, the churn behind [#59](https://github.com/datum-cloud/dns-operator/issues/59). + +**One pass over a zone:** + +1. For each record set of the zone that is not being deleted, add the finalizer and the owner reference, + and set `Accepted`. +2. Pick the holder of each name and record type, with the same function the webhook uses. The function + reads the tenant object's creation time, which the replicator copies onto the downstream object. A + rebuilt downstream cluster recreates every object at once, and ordering by those new times would fall + back to names and could move a name to another claimant. +3. Send the holders' records to PowerDNS in one call. +4. Set each record set's status for each name it lists. A record set that is not the holder gets + `Programmed=False` with reason `NotOwner`. A status is patched only when it changed, so an unchanged + zone makes no writes. +5. Remove the finalizer of each record set that is being deleted, after the write without its names has + succeeded. + +**The call to PowerDNS** replaces `EnsureRecordSet` and `DeleteRecordSet`. + +- It reads the zone once. +- It writes a name only when PowerDNS does not already hold the holder's current records. +- It deletes a name that the agent owns and that no record set claims. The agent owns a name when the + name carries its ownership note, or when a record set that is being deleted declares it. +- A pass that would delete more than 30% of a zone holding at least 10 records deletes nothing. It sets a + condition of its own on the `DNSZone` instead, which only this reconcile writes. These are octoDNS's + defaults (`MAX_SAFE_DELETE_PCENT` and `MIN_EXISTING_RECORDS` in + [`plan.py`](https://github.com/octodns/octodns/blob/main/octodns/provider/plan.py)), and we keep them + until the dry-run counts below show a reason to differ. +- PowerDNS refuses a whole request when it cannot store one record in it. A refused request is split in + halves until the refused record is alone, so one record set's bad content fails only its own names. + Halving costs a few requests where retrying name by name would cost one per name, and each accepted + request raises the zone's SOA serial. +- Any other failure is returned, and the zone is retried with backoff. + +## What users and on-call see + +- A record set that loses a name reports `Programmed=False` with reason `NotOwner` on that name, as it did + before `v0.7.0`. So the `DNSRecordRejected` alert, which fires on such a reason, can fire again + ([runbook](https://github.com/datum-cloud/infra/blob/main/docs/runbooks/dns/record-rejected.md)). +- When a claimant is deleted or drops a name, the name answers with the next claimant's records, instead + of going silent. +- A pass stopped by the deletion limit shows as a condition on the `DNSZone`, and nothing is deleted + until someone looks. + +## Phased rollout + +Each step is a pull request and a release. Before it merges, it passes the unit, envtest and PowerDNS +integration tests and every Chainsaw suite on three kind clusters. It then runs on staging, watched for +reconcile errors, records programmed, and the drift and `DNSRecordRejected` alerts, before it reaches +production. + +1. **One election function**, used by the webhook. No change in behaviour. +2. **The tenant's creation time on each downstream object**, copied by the replicator. No change in + behaviour until step 4 reads it. +3. **The call to PowerDNS**, added with its integration tests, not yet used. No change in behaviour. +4. **The per-zone reconcile.** The records served change only where a name is claimed twice, and where + leftovers are removed. A new end-to-end scenario, `claimed-names`, fails before this step and passes + after it. +5. **Removal of what step 4 made unused:** the per-record-set calls, the old reconciler + ([#86](https://github.com/datum-cloud/dns-operator/issues/86)), and the ownership parameters on the + backend interface ([#88](https://github.com/datum-cloud/dns-operator/issues/88)). + +Before step 4 reaches staging or production, a dry run of the pass counts, in that environment, the names +claimed twice and the leftovers it would delete. Names claimed twice in production are removed by hand +first, with their owners. + +## Risks + +- **Scale.** Every change to a record set queues a full pass over its zone, so the largest zone sets the + cost. Step 4 is load-tested locally with at least as many record sets as production's largest zone, + and twice its largest burst of new ones, each programmed within the 60-second target of [#59](https://github.com/datum-cloud/dns-operator/issues/59). It then + runs on staging's largest zone. On 2026-09-29 the largest production zone, which holds gateway + addresses, had 2,399 record sets, its largest burst in 30 days was 100 new ones in five minutes, and + staging's largest zone had 1,556. +- **Names claimed twice at rollout.** Step 4 brings `NotOwner`, and with it `DNSRecordRejected`, back. A + pair left in place pages on-call, and its name moves to the older claimant. +- **Leftovers.** The first pass in each zone deletes records that carry an ownership note but that no + record set claims. The dry-run count shows how many before the step ships, and the deletion limit stops + a pass that would delete more. +- **Configuration.** The agent decodes its server config strictly, so it refuses to start on a key it + does not know. Step 5 removes a key only after no deployed config sets it. From 006eb2dcff34886bd87a1e1e9e1372e6d33e3bf7 Mon Sep 17 00:00:00 2001 From: Viktor Voytovych Date: Tue, 29 Sep 2026 17:28:10 +0300 Subject: [PATCH 2/5] docs: Record the per-zone writer's local load test A local load test of the prototype met the 60-second target at production size. It also showed that a pass which writes one record set's status after another waits on the API server once per record set, so a pass now sends its writes in parallel, at most 16 at once. --- docs/enhancements/dnsrecordset-per-zone-writer.md | 9 ++++++++- 1 file changed, 8 insertions(+), 1 deletion(-) diff --git a/docs/enhancements/dnsrecordset-per-zone-writer.md b/docs/enhancements/dnsrecordset-per-zone-writer.md index 464f73b..404ea3b 100644 --- a/docs/enhancements/dnsrecordset-per-zone-writer.md +++ b/docs/enhancements/dnsrecordset-per-zone-writer.md @@ -112,6 +112,9 @@ a reconcile that queued its own status writes would run in a loop, the churn beh 5. Remove the finalizer of each record set that is being deleted, after the write without its names has succeeded. +Steps 1, 4 and 5 write to many record sets in a large zone. The pass sends those writes in parallel, at +most 16 at once, so it does not wait on the API server once per record set. + **The call to PowerDNS** replaces `EnsureRecordSet` and `DeleteRecordSet`. - It reads the zone once. @@ -168,7 +171,11 @@ first, with their owners. and twice its largest burst of new ones, each programmed within the 60-second target of [#59](https://github.com/datum-cloud/dns-operator/issues/59). It then runs on staging's largest zone. On 2026-09-29 the largest production zone, which holds gateway addresses, had 2,399 record sets, its largest burst in 30 days was 100 new ones in five minutes, and - staging's largest zone had 1,556. + staging's largest zone had 1,556. A local test of the prototype on 2026-09-29 met the target: with + 2,500 record sets in one zone, the last of 200 new ones was programmed within 2 seconds, and within + about 21 seconds with 25 ms added to each API write. The same test is why the pass writes in + parallel: one at a time, 2,500 new record sets took 115 seconds with 10 ms added to each write, and 8 + seconds with 16 at once. - **Names claimed twice at rollout.** Step 4 brings `NotOwner`, and with it `DNSRecordRejected`, back. A pair left in place pages on-call, and its name moves to the older claimant. - **Leftovers.** The first pass in each zone deletes records that carry an ownership note but that no From 39ce349d711a82c7dc79a1f90dae3db3d74174a8 Mon Sep 17 00:00:00 2001 From: Viktor Voytovych Date: Tue, 29 Sep 2026 18:21:58 +0300 Subject: [PATCH 3/5] docs: Correct who writes a zone's SOA and NS records The zone controller writes a zone's SOA and NS records only when it creates the zone. After that they are record sets like any other, so the per-zone pass writes them. The serial the SOA carries stays a separate change (#101). --- docs/enhancements/dnsrecordset-per-zone-writer.md | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/docs/enhancements/dnsrecordset-per-zone-writer.md b/docs/enhancements/dnsrecordset-per-zone-writer.md index 404ea3b..f74b918 100644 --- a/docs/enhancements/dnsrecordset-per-zone-writer.md +++ b/docs/enhancements/dnsrecordset-per-zone-writer.md @@ -61,8 +61,9 @@ still has to decide what gets through, as a Gateway API controller does for conf first. - Replacing the ownership notes. Until a zone is fully declared, they are the only mark of which records the agent may delete. -- The SOA serial, and the zone controller's writes of the SOA and NS records. Moving those into this - writer is a separate decision. +- Who sets the SOA serial. A zone's SOA and NS records are record sets, which the pass writes like any + other; the zone controller writes them only when it creates the zone. The serial the SOA carries is a + separate change ([#101](https://github.com/datum-cloud/dns-operator/issues/101)). ## Background From f977b283699a42c387e5a133ac6e8b7d97550fbe Mon Sep 17 00:00:00 2001 From: Viktor Voytovych Date: Fri, 2 Oct 2026 20:09:32 +0300 Subject: [PATCH 4/5] docs: Keep each zone's pass to its own project's records The replicator keeps one zone per domain with its accounting ConfigMap, and nothing else writes that ConfigMap. A zone made another way has none, so a project's zone for the same domain is not refused, and EnsureZone treats the PowerDNS zone that already exists as its own. A pass therefore deletes only records that a record set of its own project wrote, and the dry run before rollout counts any domain that more than one project holds. --- docs/enhancements/dnsrecordset-per-zone-writer.md | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/docs/enhancements/dnsrecordset-per-zone-writer.md b/docs/enhancements/dnsrecordset-per-zone-writer.md index f74b918..580f487 100644 --- a/docs/enhancements/dnsrecordset-per-zone-writer.md +++ b/docs/enhancements/dnsrecordset-per-zone-writer.md @@ -121,7 +121,9 @@ most 16 at once, so it does not wait on the API server once per record set. - It reads the zone once. - It writes a name only when PowerDNS does not already hold the holder's current records. - It deletes a name that the agent owns and that no record set claims. The agent owns a name when the - name carries its ownership note, or when a record set that is being deleted declares it. + name carries an ownership note from a record set of the zone's own project, or when such a record set + that is being deleted declares it. So a pass never deletes a record another project wrote, even if two + projects ever hold one domain. - A pass that would delete more than 30% of a zone holding at least 10 records deletes nothing. It sets a condition of its own on the `DNSZone` instead, which only this reconcile writes. These are octoDNS's defaults (`MAX_SAFE_DELETE_PCENT` and `MIN_EXISTING_RECORDS` in @@ -162,8 +164,8 @@ production. backend interface ([#88](https://github.com/datum-cloud/dns-operator/issues/88)). Before step 4 reaches staging or production, a dry run of the pass counts, in that environment, the names -claimed twice and the leftovers it would delete. Names claimed twice in production are removed by hand -first, with their owners. +claimed twice, the domains that more than one project holds, and the leftovers it would delete. Names +claimed twice in production are removed by hand first, with their owners. ## Risks From f6c1c5a208ebb50124be56886208005c8d21179f Mon Sep 17 00:00:00 2001 From: Viktor Voytovych Date: Mon, 5 Oct 2026 18:52:16 +0300 Subject: [PATCH 5/5] docs: Compare records by meaning in the per-zone writer PowerDNS stores some values in another form, so a text comparison would rewrite those names on every pass and raise the zone's serial each time. --- docs/enhancements/dnsrecordset-per-zone-writer.md | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/docs/enhancements/dnsrecordset-per-zone-writer.md b/docs/enhancements/dnsrecordset-per-zone-writer.md index 580f487..ecd5c44 100644 --- a/docs/enhancements/dnsrecordset-per-zone-writer.md +++ b/docs/enhancements/dnsrecordset-per-zone-writer.md @@ -120,6 +120,11 @@ most 16 at once, so it does not wait on the API server once per record set. - It reads the zone once. - It writes a name only when PowerDNS does not already hold the holder's current records. +- It compares records by meaning, not by text, because PowerDNS stores some values in another form. For + example, `2001:0DB8:0000:0000:0000:0000:0000:0001` comes back as `2001:db8::1`. The same happens to the + zone's name inside a `CNAME` or `MX` target, to `TLSA` hex, and to a `TXT` value that PowerDNS unescapes or + splits into 255-byte strings. A text comparison would rewrite each such name on every pass, and every + write raises the zone's SOA serial. - It deletes a name that the agent owns and that no record set claims. The agent owns a name when the name carries an ownership note from a record set of the zone's own project, or when such a record set that is being deleted declares it. So a pass never deletes a record another project wrote, even if two