Skip to content

Respect cgroup limits - #38

Open
tmakatos wants to merge 1 commit into
add_psi_monitoringfrom
respect_cgroup_limits
Open

tmakatos wants to merge 1 commit into
add_psi_monitoringfrom
respect_cgroup_limits

Conversation

@tmakatos

Copy link
Copy Markdown
Collaborator

A VM process can have less CPU capacity than its vCPU count and
host-wide utilisation imply. When its cgroup is being throttled, adding
another IO thread cannot create CPU capacity and can instead drive
repeated, counterproductive scale-up decisions.

Resolve each process's cgroup v2 path, read its cpu.max quota and
cpu.stat throttle counters, and retain the latest sample on the
instance. The threshold engine can then hold scale-up whenever
throttled_usec increased since the preceding sample.

The first refresh always establishes an initial sample. The optional
refresh_cgroup_on_each_read setting refreshes quota and counters on
every tick when operators need runtime cgroup changes reflected
immediately; leaving it disabled avoids repeated filesystem reads for
stable production placement. Component-test configuration explicitly
disables both refresh and cgroup-based scale-up blocking where the fake
process does not model cgroups.

Signed-off-by: Thanos Makatos thanos.makatos@nutanix.com


Stack created with GitHub Stacks CLI • Give Feedback 💬

@tmakatos
tmakatos added this pull request to stack #43 September 24, 2026 21:44
@tmakatos
tmakatos removed this pull request from stack #43 October 1, 2026 11:52
@tmakatos
tmakatos added this pull request to stack #71 October 1, 2026 12:18
@tmakatos
tmakatos requested a review from lforchini as a code owner October 1, 2026 16:23
@tmakatos
tmakatos force-pushed the respect_cgroup_limits branch from 49e82e1 to 0a909a8 Compare October 1, 2026 16:23
@tmakatos
tmakatos force-pushed the respect_cgroup_limits branch from 0a909a8 to 4dc6138 Compare October 1, 2026 16:35
@tmakatos
tmakatos force-pushed the respect_cgroup_limits branch from 4dc6138 to f738f6a Compare October 1, 2026 16:42
A VM process can have less CPU capacity than its vCPU count and
host-wide utilisation imply. When its cgroup is being throttled, adding
another IO thread cannot create CPU capacity and can instead drive
repeated, counterproductive scale-up decisions.

Resolve each process's cgroup v2 path, read its `cpu.max` quota and
`cpu.stat` throttle counters, and retain the latest sample on the
instance. The threshold engine can then hold scale-up whenever
`throttled_usec` increased since the preceding sample.

The first refresh always establishes an initial sample. The optional
`refresh_cgroup_on_each_read` setting refreshes quota and counters on
every tick when operators need runtime cgroup changes reflected
immediately; leaving it disabled avoids repeated filesystem reads for
stable production placement. Component-test configuration explicitly
disables both refresh and cgroup-based scale-up blocking where the fake
process does not model cgroups.

Signed-off-by: Thanos Makatos <thanos.makatos@nutanix.com>
@tmakatos
tmakatos force-pushed the respect_cgroup_limits branch from f738f6a to 03b6139 Compare October 2, 2026 05:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant