When configuring an Azure Virtual Machine Scale Set (VMSS), you have the option to define autoscaling rules. Autoscaling is the process of dynamically allocating resources to match performance requirements. As the volume of work grows, an application may need additional resources to maintain the desired performance levels and meet service level agreements (SLAs). When demand decreases and those additional resources are no longer required, they can be automatically removed to minimise costs.
As part of the autoscale configuration within an Azure Virtual Machine Scale Set, we can set both a duration and a cool down period. In this post, I will explain the differences between these two options using a couple of practical scenarios, and I’ll also walk through a simple experiment to show how the cool down duration behaves in action.
Note that this demo reflects what I observed during my own testing. The behaviour may differ from how some explanations describe the cool down period.
Want to learn more about scaling in Azure?
If you’d like to dive deeper into Azure VM Scale Sets, you can start with the Microsoft Learn article, Azure Virtual Machine Scale Sets Overview.
In addition to Azure VM Scale Sets, several other Azure services support autoscaling, including Azure App Service Plans. With App Service Plans, you can configure autoscale rules based on metrics such as CPU usage, memory usage, HTTP queue length, and more. Essentially, the App Service Plan includes a built‑in VM Scale Set under the hood. Learn more: Azure App Service Plans
For autoscaling best practices, and to explore the range of Azure services that include built‑in scaling capabilities, visit the Microsoft Learn article, Autoscaling guidance – Best practices for cloud applications.
Note: It’s important to plan and configure your scaling rules carefully to avoid performance issues, unnecessary scaling, or unexpected costs caused by incorrect settings. The metrics used in this post are purely for demonstration purposes to observe how the cool down period works.
Let’s get started
Duration in Azure VM Scale Set
Duration is the time window the VM Scale Set will look back over when evaluating metrics before deciding whether to scale.
For example, in the scaling rule below, I have configured:
- If CPU Percentage is greater than 85%
- for a DURATION of 10 minutes
- Increase the Instance/VM (Virtual Machine) count by 1
For this condition to trigger, the CPU must remain above 85% continuously for the full 10 minute duration. The VM Scale Set reviews CPU utilisation over the past 10 minutes, and if the CPU has consistently been above 85%, it will add another instance/VM.
Below is a screenshot of an Azure VM Scale Set rule. I have used a green arrow to highlight the DURATION field.

Cool down in Azure VM Scale Set
The COOL DOWN period comes into effect after a scale‑in (removing a VM instance) or scale‑out (adding a VM instance) event is triggered. For example, if I set a COOL DOWN period of 10 minutes, this instructs the scale set not to scale again for another 10 minutes. Essentially, you’re asking the scale set to pause during that COOL DOWN window so the environment can stabilise and the system can assess whether the previous scaling action has had an impact on CPU utilisation.
I have used a blue arrow to highlight the COOL DOWN field in the image below.

Let’s take a look at how the scale rule behaves on a timeline.
Note: Based on the demos you’ll see shortly, I found that the VM Scale Set evaluated the full 10 minute duration using all available metrics, including those collected during the cooldown period. Continue to learn more.
In the diagram below, we start with one VM in our scale set.

1. At 3pm CPU for VM1 goes above the threshold of 85%, and constantly remains above the threshold for a duration of 10 minutes (until 3.10pm).
2. At 3:10pm, the scale set looks back over the previous 10 minutes (from 3:00pm to 3:10pm). Because the CPU was consistently above 85% for that entire window, the scale set adds another VM, bringing the total to two VMs.
3. At the same moment (3:10pm), the cool down period begins. During this cool down window, no additional scale‑in or scale‑out actions will occur. However, based on my demo, the cool down period did not pause time or stop metric collection. As shown in the diagram, adding a second VM at 3:10pm reduces CPU utilisation to around 50–60%, but metrics continue to be gathered and evaluated in the background. The cool down period simply instructs the scale set to temporarily pause scaling actions, not to stop analysing data.
4. The additional VM stabilises the workload, and the scale set continues operating normally with CPU averaging between 50% and 60%.
5. At 3:40pm, CPU utilisation rises again and remains above 85% for another full 10 minute duration.
At 3:50pm, the scale set looks back over that 10 minute window and decides to add another VM.
I then moved onto another scenario to see how the scale set responds under different conditions.
What if, after the second VM was added, the CPU did not stabilise and instead remained consistently above 85%? Would the VM Scale Set add another VM immediately after the cooldown period ends, or would it wait for another full 10‑minute duration before deciding to scale out again?
Note: This is exactly why it’s important to plan your scaling rules carefully. A reminder that this post was created to test a few scenarios and observe how the VM Scale Set duration and cool down period behave under different conditions.
Firstly, if you come across a scenario where the VMSS adds an additional VM and it still doesn’t make a difference, for example, the CPU remains above 85%, then you should consider investigating and possibly reconfiguring your scaling rule. That usually indicates the rule isn’t aligned with the workload pattern.
However, Because I am experimenting, I want to understand what happens next.
To test this behaviour, I deployed a new Virtual Machine Scale Set and configured a VM Scale Set rule as follows:
- If CPU Percentage is greater than 0.1% (Yes, I know a metric you would never use in production, but it’s for testing purposes only!)
- for a DURATION of 10 minutes
- Increase the VM (Virtual Machine) count by 1
- COOL DOWN period of 10 minutes

The diagram above shows that I start with one Virtual Machine at 3pm. Because of the intentionally low CPU threshold of 0.1%, the condition is met instantly. The VM Scale Set then looks back over the previous 10‑minute duration (from 3:00pm to 3:10pm) and adds another VM instance. A 10 minute cooldown is triggered to pause any further scaling actions while the scale set stabilises.
However, as shown in the diagram, it was interesting to see that metrics continue to be monitored and recorded in the background. The additional VM has not made any difference, as CPU utilisation remains above 0.1%.
At 3.20pm, the cooldown period of 10 minutes expires.
But what happens now?
CPU has remained above 0.1% throughout the entire cooldown period, and the extra VM has not reduced utilisation. So what does the VM Scale Set do next?
- Does it immediately trigger another scale‑out event and add a third VM right after the cooldown ends at point A (3:20pm)?
- Or does it wait for another full 10‑minute duration before scaling again, adding the next VM at point B (3:30pm)?
My findings from the demo showed the following:
Answer:
Another VM was added shortly after the cooldown period ends at around 3:20pm. The VM Scale Set takes into account the previous 10 minutes of metrics, including those collected during the cooldown period. The cooldown period was used to pause scaling actions, it did not pause time or stop metric collection. Under the hood, metrics were still being analysed and recorded continuously.
Below are the scaling logs from my testing:

Let’s zoom in and focus on the timestamp column, which shows the exact times the VM Scale Set added a new VM instance. See the image below.

According to the time stamps above:
- at 9.07am a VM scale-out operation was triggered to add another VM/Instance. This happened because CPU utilisation was greater than 0.1% for a 10 minute duration. The VM Scale Set looked back over the window from approximately 8:57am to 9:07am and confirmed the condition was met.
- When the scale‑out event occurred, a 10 minute cooldown was also initiated to allow the new VM to deploy and for CPU utilisation to stabilise.
- The cooldown period ended at around 9:17am. However, immediately after the cooldown expired, another scale out operation was triggered. The VM Scale Set did not wait for another 10 minute duration.
- Because the CPU threshold was set to 0.1%, the VM Scale Set never stabilised. CPU remained above the threshold continuously, so the system kept scaling out, adding another VM after each cooldown period.
What did I find?
When the VM Scale Set looks back at the duration window, it does include the metrics collected during the cooldown period. The cooldown only pauses scaling actions, it does not pause metric collection. Metrics continue to be gathered and analysed in the background, and those metrics are used to decide whether another scale operation is required.
Note: Don’t forget to configure a scale‑in rule so the scale set can remove VMs when CPU levels drop during quieter periods. Without a scale‑in rule, the VM Scale Set will only ever scale out to the maximum number of VMs you’ve configured, as it has no way of knowing when to scale back in.
In my demo, because the CPU threshold was set to 0.1%, the VM Scale Set would never stabilise and therefore would never scale in. Of course, this is something you would never do in production; the purpose of this demo was purely to observe how the cooldown and duration periods behave.
Always plan and configure your VM Scale Set rules according to your workload requirements.
And that’s it for now. I hope you found this post useful, and if you have any feedback, please feel free to leave a comment below.
See you in the next post.


Hi,
Thank you, very interresting.
Could you please tell us, make a demo like this one, with the same example but with a Duration of 15 minutes.
To see if action is trigger just after the cooldown period or not?
Thanks