Skip to content

Validate scale actions using IOPS instead of cumulative IOs - #70

Open
tmakatos wants to merge 1 commit into
print_start_msg_firstfrom
fix_validation
Open

tmakatos wants to merge 1 commit into
print_start_msg_firstfrom
fix_validation

Conversation

@tmakatos

@tmakatos tmakatos commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

Post-scale validation compared lifetime I/O totals as if they were
rates. Example: baseline 165453639 ops, observed 171620504 (+3.7%) after
the validation polls, required 173726321 (+5% of the lifetime total), a
steady multi-hundred-kops workload still failed validation and was
reverted.

Scale-up is still proposed when CPU utilisation is above the scale up
threshold; validation is meant to require at least scale_up_min_gain
more IOPS afterward. Applying that percent to lifetime totals made even
healthy post-scale rates look like failures.

Compare IOPS over the pre-scale baseline window (from the last
settled baseline reset through the scale decision; not a dedicated
duration knob) to IOPS over the post-scale window governed by
scale_validation_sample_polls.

When validating a scale down that ran because CPU utilisation was low,
do not revert just because IOPS fell with demand: if utilisation is
still at or below the scale up threshold, keep the smaller pool.

Enrich the existing scale up decision log with current_iops,
min_target_iops, and validation_after_secs to make it clear what the
minimum performance must be achieved and when for the scale up to be
considered successful.

AI-generated: everything, manual fixup of commit message
Signed-off-by: Thanos Makatos thanos.makatos@nutanix.com


Stack created with GitHub Stacks CLI • Give Feedback 💬

@tmakatos
tmakatos added this pull request to stack #71 October 1, 2026 12:18
@tmakatos
tmakatos requested a review from lforchini as a code owner October 1, 2026 16:23
@tmakatos
tmakatos force-pushed the fix_validation branch 2 times, most recently from 0f99df5 to c87327b Compare October 2, 2026 05:05
Post-scale validation compared lifetime I/O totals as if they were
rates. Example: baseline 165453639 ops, observed 171620504 (+3.7%) after
the validation polls, required 173726321 (+5% of the lifetime total), a
steady multi-hundred-kops workload still failed validation and was
reverted.

Scale-up is still proposed when CPU utilisation is above the scale up
threshold; validation is meant to require at least scale_up_min_gain
more IOPS afterward. Applying that percent to lifetime totals made even
healthy post-scale rates look like failures.

Compare IOPS over the pre-scale baseline window (from the last
settled baseline reset through the scale decision; not a dedicated
duration knob) to IOPS over the post-scale window governed by
scale_validation_sample_polls.

When validating a scale down that ran because CPU utilisation was low,
do not revert just because IOPS fell with demand: if utilisation is
still at or below the scale up threshold, keep the smaller pool.

Enrich the existing scale up decision log with current_iops,
min_target_iops, and validation_after_secs to make it clear what the
minimum performance must be achieved and when for the scale up to be
considered successful.

AI-generated: everything, manual fixup of commit message
Signed-off-by: Thanos Makatos <thanos.makatos@nutanix.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant