Skip to content

cluster DRS: power off under-utilized hosts and wake them on demand - #14371

Open
rootnomad wants to merge 8 commits into
apache:mainfrom
rootnomad:pr/drs-power-management
Open

rootnomad wants to merge 8 commits into
apache:mainfrom
rootnomad:pr/drs-power-management

Conversation

@rootnomad

@rootnomad rootnomad commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Description

Extends the existing cluster DRS service with optional host power management. When a cluster
drops below a configurable low utilization threshold, DRS disables an empty, out-of-band
manageable host, records a durable host detail marking that DRS powered it off, and issues the
out-of-band power-off. When the cluster rises above a high threshold, DRS powers a host it turned
off earlier back on. At most one host is powered off or on per poll, and the remaining hosts must
be able to carry the load before a host is released.

A host is only powered off once it is genuinely empty. If no host is empty, an optional active
evacuation mode drains the least-loaded releasable host (migrating its VMs to the others within
the DRS migration budget) so a later poll can power it off once empty. Draining is disabled first
so the allocator stops targeting it, and a drain that stops making progress (its migrations keep
being rejected for affinity, host tags or storage) is abandoned and the host returned to service
rather than being re-submitted forever. An optional predictive mode folds a short least-squares
forecast of recent utilization into the decision, so a rising trend wakes a host sooner and a
transient dip does not trigger a power-off.

The power-off path is ordered so the durable marker is written before the irreversible power
operation: a management-server crash in between leaves a host DRS still recognises and powers back
on, never one stranded off. A host that comes back by any path is reconciled (re-enabled, unmarked)
on the next poll, and a power-off whose command reports failure is left disabled and marked for
reconciliation rather than assumed still running.

Everything is off by default behind per-cluster settings: drs.power.management.enable,
drs.power.management.low.threshold, drs.power.management.high.threshold,
drs.power.management.evacuate.enable, drs.predictive.enable and drs.predictive.window. It requires
automatic DRS and out-of-band management to be configured on the cluster. No schema change: it uses
a host detail and config keys.

Types of changes

  • Breaking change (fix or feature that would cause existing functionality to change)
  • New feature (non-breaking change which adds functionality)
  • Bug fix (non-breaking change which fixes an issue)
  • Enhancement (improves an existing feature and functionality)
  • Cleanup (Code refactoring and cleanup, that may add test cases)
  • Build/CI
  • Test (unit or integration test code)

Feature/Enhancement Scale or Bug Severity

Feature/Enhancement Scale

  • Major
  • Minor

How Has This Been Tested?

Unit tests cover the decision and lifecycle logic (38 tests across ClusterDrsPowerManagementTest
and ClusterDrsPowerOrchestrationTest): the release and wake thresholds, the predictive forecast,
the first-fit placement pre-check, the power-off ordering (marker before the power operation,
disable gated on the state transition, no rollback on an unconfirmed failure), the power-on marker
handling, the in-flight-plan guard, and the drain lifecycle (wait while a batch migrates, abandon
on no progress or when VMs no longer fit, disable the candidate, power the emptied host off only
while it is still manageable and releasable).

Verified end to end against a BMC wired to the hypervisor's VM power control, so an IPMI power
command actually stops or starts the host. On a management server running this build, with a
cluster whose host had out-of-band management configured against that BMC and power management
enabled:

  • Power off: with the cluster idle, DRS powered off the empty out-of-band-manageable host. The
    management log recorded "disabling and powering off empty host", the host moved to the Disabled
    resource state with a drs.power.state=off host detail, its agent was detached without
    investigation, and the IPMI power-off actually stopped the host VM, which went down.
  • Wake: with the cluster then driven above the high threshold, DRS powered the same host back on.
    The log recorded "powering on host" and the IPMI power-on actually started the host VM. Once it
    booted and its agent reconnected, DRS re-enabled the host and cleared the marker on the next
    poll.

Power management was also observed to correctly skip the cluster while a DRS migration plan was in
flight.

How did you try to break this feature and the system with this change?

The logic was audited for the failure paths that matter for powering real hardware off: a
management-server crash in the window between disabling, marking and powering off a host; a
power-off command that reports failure after the host already went down; an accepted but
unconfirmed power-on; a host returning by a path other than a DRS wake; the allocator placing new
VMs on a host mid-drain; a drain whose VMs cannot actually be migrated; and a host going down or
being disabled by an operator mid-drain. Power management is skipped entirely for a cluster that
has a DRS migration plan in flight so a migration source or destination is never powered off under
it. The sub-minute allocation race between powering a host off and its agent disconnecting is
handled by disabling the host first and detaching its agent without investigation.

…ation

Extend the cluster DRS poll to power manage hosts. When a cluster's load falls
below a low threshold, DRS releases an empty, out-of-band capable host: it
disables the host first so no new workload is scheduled onto it, then powers it
off through the out-of-band management driver, rolling the resource state back
if the power operation fails. When load rises above a high threshold, a host
that DRS previously powered off is woken through the same driver and re-enabled.

Add an opt-in predictive mode that forecasts short-term cluster utilisation
from a rolling history with a least-squares trend, so a host can be woken ahead
of rising demand rather than after it.

Add an opt-in active evacuation mode that migrates running VMs off a releasable
host, capped by the existing per-run migration limit and aware of reserved
capacity so it never pushes a destination host past its headroom.

All of this is off by default and gated by ConfigKeys. Unit tests cover the
release and wakeup classification, the disable-then-power-off path with
rollback, the utilisation forecast and the capped evacuation planner.
Fixes found by auditing the power orchestration:
- persist the wake marker before the irreversible power-off and remove it
  on rollback, so a crash between the power-off and the marker write no
  longer strands a host powered off and unrecognised.
- abort the power-off when the host could not be disabled, instead of
  powering it off while CloudStack still believes it is schedulable.
- keep the marker and leave the host disabled after issuing power-on, and
  reclaim it (re-enable, clear the marker) only once the host is actually
  Up, so an accepted but unconfirmed power-on cannot strand the host, and
  a host returned by any path is never left disabled and marked forever.
- skip power management for a cluster with a DRS migration plan in flight,
  so a migration source or destination is never powered off under it.
- cap each evacuation destination at the high utilization threshold so
  draining a host does not push another host past it.

Adds unit tests for the power-off ordering, disable gating and rollback,
the power-on marker handling, and the in-flight-plan guard.
…ress

The evacuation path drained the least-loaded releasable host by submitting
migrations and re-picking it each poll until empty. A VM that is unplaceable
for affinity, host tags or storage passes the capacity pre-check but is
rejected by the migration itself, so the host was re-submitted every poll
and left half-drained.

Track a drain with a distinct draining host marker and the host's VM count
at the last batch. Continue a drain already in progress before starting
another, wait while a batch is still in flight, and abandon the drain
(clearing the marker) once it makes no progress across a poll or the
remaining VMs no longer fit, instead of re-submitting forever. Power-off
replaces the draining marker with the off marker once the host is empty.
…drain event

- detach the agent without investigation before cutting power, so the link
  drop from a DRS power-off is not reported as a host-down failure. The host
  is already disabled and empty and is going away intentionally.
- resolve the concrete VM and destination pairs before opening the drain
  action event, and open it only when the batch is non-empty, so a poll that
  finds nothing actually migratable no longer leaves a start event with no
  matching terminal event.
powerOffHost removed the wake marker and re-enabled the host whenever the
power-off command threw. A failure response does not prove the host is
still running: a chassis power-off can take effect and still report a
timeout or error, which left the host powered off, enabled and unmarked,
so DRS never woke it and the scheduler kept targeting a dead host. Leave
the host disabled and marked on failure; the next poll re-enables and
unmarks it if it is still up, or treats it as a DRS-powered-off host if it
is actually down.
…state

A host being drained stayed Enabled, so the allocator could keep placing new
VMs on it (it is the least-loaded host, which the allocator prefers), fighting
the drain and tripping the no-progress abandon. Draining hosts were also only
reconciled while in upHosts, so one that went Down or was disabled mid-drain
leaked its marker and progress entry.

Disable the candidate when a drain starts so the scheduler leaves it alone, and
re-enable it if the drain is abandoned. Classify draining hosts in their own
bucket: continue the drain while the host is up (regardless of whether
evacuation is still enabled, so toggling it off does not strand a host) and
abandon it if the host is no longer up. Count a draining host's load in cluster
utilization, power off a host that has finished draining directly (powerOffHost
now tolerates an already-disabled host), and only start a new drain when no host
is empty to power off.
…ompletion

- Persist the draining marker before disabling the host when a drain starts,
  mirroring the marker-before-action ordering powerOffHost uses. If job
  submission or the management server fails in the window after the disable,
  the host is still recognised as draining and reconciled on the next poll
  instead of stranded disabled with no marker. Undo the marker if the disable
  itself fails.
- When a drain completes, power the emptied host off only if it is still
  out-of-band manageable and the cluster can still release it; otherwise
  return it to service, so a host is not flapped off when management went away
  or load rose during the drain.
@rootnomad
rootnomad marked this pull request as draft October 8, 2026 18:52
@rootnomad
rootnomad marked this pull request as ready for review October 8, 2026 19:42

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant