Skip to content

Keep projects in Terminating until their resources drain - #752

Merged
scotwells merged 3 commits into
mainfrom
fix/750-namespace-drain-before-project-delete
Aug 6, 2026
Merged

scotwells merged 3 commits into
mainfrom
fix/750-namespace-drain-before-project-delete

Conversation

@scotwells

@scotwells scotwells commented Aug 6, 2026 •

Copy link
Copy Markdown
Contributor

Closes #750.

Impact

Deleting a project could strand resources that operators created for it. The project finished deleting while objects inside it still held finalizers, so those finalizers never ran.

For network-services-operator, gateway configuration outlived the project. It kept propagating to the edge and serving traffic, and NSO held a finalizer it could never release, failing reconcile forever.

A project now stays in Terminating until its control plane is empty, and reports what blocks it. Consumers finish their work against a control plane that still answers. A project with nothing stuck still deletes in seconds.

What changed

Namespaces are no longer force-finalized. Kubernetes guarantees a namespace outlives its contents, and consumers resolve state through it. Namespaces are deleted and left to the namespace controller each project control plane already runs.

The completion check covers default. It previously accepted "no namespaces left". The API server refuses to delete default, so resources held there stayed invisible and the project completed on top of them. Completion now requires default to be empty.

The ResourceCleanup condition names the blocker:

Waiting for project resources to be removed: configmaps "default/held-resource" (finalizers: test.miloapis.com/consumer-cleanup)

Two per-project metrics back it and clear on completion: milo_resourcemanager_project_deletion_pending_seconds and milo_resourcemanager_project_deletion_blocking_resources.

The project control plane is deleted after cleanup rather than before, so it no longer disappears while a consumer is still finalizing against it.

Deletion stays asynchronous, as #553 made it. A stuck project polls on its own and holds up nobody else.

There is no escape hatch. A stuck project means a consumer bug, which should be visible and fixed rather than overridden. If operators need one later, make it an explicit gated admin action.

Testing

New e2e project-deletion-blocked-resources covers both places project resources live: held in default, and held in a project namespace. Neither project may be removed while its resource is held, and both must delete once the finalizers are released.

Against unmodified Milo it fails as expected. The default case loses its project 45 seconds after deletion while the held ConfigMap remains.

Verified end to end against network-services-operator on a two-cluster environment. The gateway finalizer runs while the namespace exists, every downstream object is removed, and the project deletes cleanly.

While getting the new test green, project-creation turned out to share the cluster-scoped User user-admin with project-deletion. The two run concurrently, so one could delete it while the other waited for it to be ready — project-creation now uses its own user.

Unit tests cover the completion check.

Note on #538

#538 fixes the field the force-finalize path clears. This PR removes that path. Force-finalizing is inert today because the guard returns early on a successful get, so repairing it would arm the behaviour #750 describes.

scotwells and others added 3 commits August 6, 2026 08:28
…ces drain

Project purge force-finalized namespaces and treated "no namespaces left" as
cleanup complete. Neither is safe for consumers: force-finalizing removes a
namespace ahead of its contents, and the completion check ignored the default
namespace entirely, so a project could finish deleting while objects in it
still carried finalizers that would now never run.

Namespaces are now deleted and left to the namespace controller wired for each
project control plane, and a project is only complete when its namespaces are
gone and the namespaces that cannot be deleted hold nothing. While resources
remain, the ResourceCleanup condition names them and the finalizers holding
them, and two metrics report how long the project has been waiting and how many
resources are holding it.

The project control plane is now torn down after cleanup completes rather than
at the start of deletion, so consumers finalize against a control plane that
still answers.

Refs #750

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VTE43NjiyoTnSn2bAzDS2o
Project creation requires organization parent information in the request, so
the test now creates its projects through the organization control plane like
the other project tests, and cleans up anything a failed run left behind so it
can be re-run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VTE43NjiyoTnSn2bAzDS2o
Both tests created and deleted the cluster-scoped User "user-admin" and run
concurrently, so one could delete it while the other was waiting for it to be
ready. project-creation now uses its own user.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VTE43NjiyoTnSn2bAzDS2o
@scotwells
scotwells marked this pull request as ready for review August 6, 2026 16:14

@mattdjenkinson mattdjenkinson left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we check if/how much this affects the deleting in the portal before a release is tagged?

@scotwells

Copy link
Copy Markdown
Contributor Author

@mattdjenkinson this does mean it'll take time to delete projects now. The time taken will depend on how many resources are in the projects.

Agree we should see what the UX in the portal is and make sure it handles situations where it may take time to delete the project.

@scotwells
scotwells merged commit 3bca885 into main Aug 6, 2026
5 checks passed
@scotwells
scotwells deleted the fix/750-namespace-drain-before-project-delete branch August 6, 2026 16:34
@scotwells

Copy link
Copy Markdown
Contributor Author

@mattdjenkinson I loaded a project with a 1000 DNS resources and issued a deletion. As expected it took ~1 minute for the deletion to complete.

The portal didn't give me any indication the project was actively deleting which was a little confusing. The portal should be checking if the deletion timestamp is set and show a little spinning indictor that it's actively being deleted. Or maybe we hide those projects by default with the option of seeing projects being deleted?

The API response does include this which we can use to show additional information to users if needed (probably unnecessary)

  - lastTransitionTime: "2026-08-07T16:58:58Z"
    message: 'Waiting for project resources to be removed: configmaps "default/billing-export-job"
      (finalizers: billing.example.com/drain-pending)'
    observedGeneration: 2
    reason: CleanupAwaitingCompletion
    status: "True"
    type: ResourceCleanup

@scotwells

Copy link
Copy Markdown
Contributor Author

@mattdjenkinson

Copy link
Copy Markdown
Contributor

Thanks for this Scot. Will get something wired up asap.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Projects can finish deleting while their resources are still being cleaned up

2 participants