A hands-on approach to test autoscaling reveals vital discrepancies between scheduler capacity and infrastructure metrics, essential for optimized operations.

The intricate dance between autoscaling and resource scheduling has gained new attention, especially as developers face surprising discrepancies between expected performance and actual operational constraints. In a recent test involving Azure Virtual Machine Scale Sets (VMSS) and Docker Swarm, we discovered a significant gap that could inform better autoscaling practices.
Initial Test Success
We initiated our autoscaling strategy by monitoring CPU usage, memory metrics, and queue lengths, structured with cooldown periods to minimize unnecessary scaling actions. This first test confirmed a successful scale-out: when CPU usage exceeded the specified threshold, Azure effectively created a new VMSS instance, which then joined the Docker Swarm, enhancing our cluster's processing capabilities.
What's critical is that during this test, the scaling metric successfully triggered the addition of resources, reinforcing confidence in the system's design. A sample command we used to establish the CPU monitoring rule was:
az monitor autoscale rule create \
–resource-group <resource-group> \
–autoscale-name swarm-autoscale \
–condition “Percentage CPU > 80 avg 10m” \
–scale out 1 \
–cooldown 5
Revealing Constraints Through Downsizing
However, when we shifted our focus to reducing worker sizes—from 16 GiB to 8 GiB—we uncovered an alarming limitation within our scaling strategy. Although we anticipated that our scheduling policy would manage the smaller nodes effectively, multiple services became incapable of successfully running due to insufficient memory availability. Strikingly, this was occurring despite Azure Monitor indicating that the VM still had adequate resources available.
This disconnect prompted a deep examination of how resource reservations functioned within the Docker Swarm setup. Scheduling a task with a specific memory reservation does not imply that the task constantly uses that much memory; it merely signifies that such capacity must be available at the time of placement.
Understanding the Divergence Between Systems
Here's the crux: the pressures perceived by Swarm as it deals with scheduling requests differ fundamentally from the metrics Azure Monitor evaluates, which focus on overall resource utilization. Hence, while Swarm struggled with task placement due to unmet memory reservations, the autoscaler remained blissfully unaware of any issues, basing its actions solely on host-level utilization metrics.
This revelation led to a critical consideration: the autoscaling algorithms need to operate in tandem with scheduler preparedness. That involves examining not just the typical host metrics but also gaining insight directly from the scheduler regarding its task placement capabilities.
Implementing Diagnostic Tools
To diagnose whether tasks remain pending due to actual runtime pressures or unmet reservation conditions, establishing a diagnostic sequence is vital. Commands like docker service ls and docker service ps <service> –no-trunc are essential tools for analyzing service statuses and potential issues.
The central question remains: is the service's inability to progress due to exhausted resources or due to the scheduler encountering constraints that it can't satisfy? The distinctions between these scenarios demand tailored approaches to scaling actions.
Rethinking Your Autoscaling Tests
As this investigation unfolded, I modified my testing approach for autoscaling systems. Instead of relying solely on successful CPU scale-out metrics, I now focus on four key assessments:
- Can all required tasks be scheduled at the minimum worker count?
- What adjustments need to be made when worker sizes change?
- Which specific metric instigates infrastructure scale-out procedures?
- Does that metric respond before the scheduler starts leaving tasks pending?
This holistic view emphasizes that policies for autoscaling and scheduling must complement each other. If one system fails to recognize the pressures felt in the other, simply amplifying autoscale commands around traditional utilization metrics won't bridge the gap effectively.
Ultimately, one of the most critical signals for autoscaling won't be the overall health of a VM based on metrics alone, but rather the indication that there are tasks waiting for execution that cannot be placed. Thus, integrating scheduler feedback into autoscaling strategies should become a priority for organizations looking to optimize operations.
This journey into the complexities of autoscaling reveals an underlying truth: fine-tuning capacity planning requires not only awareness of resource availability but also an understanding of how task scheduling interacts with those resources, a factor often overlooked in traditional monitoring and scaling practices.
Discussion
Sign in to join the discussion.