Azure App Service自动缩放CPU平均值计算异常问题问询
Is using 0.0 to replace missing metric values expected behavior?
First off, this behavior is not expected and is a known pitfall with Azure Autoscale's metric evaluation logic. Let’s break down your logs to confirm:
Looking at your MetricEvaluation log entry, the CPU data array is:
[47.666666666666664,45.428571428571431,48.666666666666664,44.0,0.0,0.0]
When calculating the average for the 5-minute window, the Autoscale service included those two 0.0 values (which correspond to missing data points) in the computation:
- Sum of all values: ~185.77
- Average: 185.77 / 6 ≈ 30.96, which exactly matches the
observedValuein yourScaleRuleEvaluationlog.
Since you confirmed CPU never dropped below 40% during this period, those 0.0 values are clearly replacements for missing metric data. This is not an intended design choice—it’s a flaw that can skew average calculations and trigger incorrect scale-in actions.
Better Temporary Workarounds (Beyond Adjusting Time Windows)
You’ve already adjusted scale-out/in time windows, but here are more targeted fixes to mitigate this issue:
Switch to Maximum CPU statistic for scale-in rules
Instead of relying on average CPU for scale-in, configure your rule to use the Maximum CPU value. Missing data points filled with 0.0 won’t affect the maximum calculation, so you’ll get an accurate picture of peak CPU usage during the window. For example, update your scale-in rule to: "If past 5 minutes maximum CPU ≤ 40% then remove 1 instance".Tighten scale-in thresholds and extend cooldown periods
Increase the scale-in cooldown from 10 minutes to 20+ minutes to ensure the rule only triggers on sustained low load, not transient metric errors. Also, lower your scale-in thresholds (e.g., change average CPU ≤50% to ≤30%)—this makes it harder for skewed averages (from 0.0 values) to meet the scale-in criteria, while still allowing scale-in when load is genuinely low.Validate metric data integrity
Check your App Service Plan’s CPU metrics in Azure Monitor to identify when/why data is missing. This could be due to temporary platform-side metric collection delays or resource constraints. If you spot consistent gaps, open a support ticket with Azure to investigate the root cause of the missing metrics.Use custom metrics for autoscale
Deploy a custom metric collector (e.g., using Azure Monitor Application Insights or a custom script) to report CPU usage directly. Custom metrics give you control over how missing data is handled (you can choose to ignore gaps instead of filling with 0.0), making your autoscale logic more reliable.
内容的提问来源于stack exchange,提问作者samfromlv

