Azure Alert-v1低CPU利用率警报不触发问题求助
Azure Alert-v1警报无法触发排查求助
我使用以下自定义Kusto查询筛选运行时长超过3天且CPU利用率较低的虚拟机,查询单独执行时可正常返回结果,但基于此创建Azure Alert-v1警报后,警报始终无法触发。附上查询语句及警报配置截图说明,恳请协助排查问题。
let startdate = ago(7d); let enddate = now(); let livemachine = Heartbeat | where TimeGenerated > ago(7d) | summarize AggregatedValue = count() by Computer | extend Uptime = AggregatedValue / 1440 | where Uptime > 3 | project Computer ; let cpu = InsightsMetrics | where Computer in (livemachine) | where Namespace contains "Processor" and Name contains "UtilizationPercentage" | where TimeGenerated between (startdate .. enddate) | project startdate, enddate, TimeGenerated, Computer, CPUUtilization = Val, SubscriptionID = _SubscriptionId, ResourceID = _ResourceId | extend ResourceGroup = split(ResourceID,"/")[4]; cpu | summarize AggregatedValue = avg(CPUUtilization) by bin(TimeGenerated, 7d), Computer, tostring(ResourceGroup), SubscriptionID, ResourceID, startdate, enddate | join kind=inner (Heartbeat | where TimeGenerated > ago(7d) | summarize AggregatedValue = count() by Computer | extend Uptime = AggregatedValue / 1440 | where Uptime > 3 | project Computer, Uptime ) on Computer | where AggregatedValue < 0.6 | project startdate, enddate, Computer, ResourceGroup, round(AggregatedValue, 2), SubscriptionID, ResourceID, TimeGenerated, Uptime | summarize arg_max(TimeGenerated, *) by Computer
警报配置截图说明
- 截图1:展示警报规则的基础配置,包含名称、描述、资源范围等内容
- 截图2:展示信号逻辑配置,包含查询关联、周期、频率、触发条件等参数
- 截图3:展示操作组相关配置内容
排查方向及解决方法
1. 硬编码时间范围与警报动态窗口冲突
查询中硬编码了startdate = ago(7d)和enddate = now(),但Azure警报会自动注入$StartTime和$EndTime参数,对应警报配置的周期和频率时间窗口。硬编码的时间会覆盖动态窗口,导致警报执行时的查询范围与手动执行不一致。
- 解决方法:替换为警报内置参数:
同时确保警报配置的周期设置为7天,频率按需调整(比如每天执行一次)。let startdate = $StartTime; let enddate = $EndTime;
2. 聚合逻辑与警报触发条件不匹配
查询最后用arg_max(TimeGenerated, *) by Computer仅保留每个虚拟机的最新记录,若警报触发条件依赖特定字段的聚合逻辑,可能无法识别有效触发项。
- 解决方法:检查警报触发条件配置:
- 若选择“基于度量值”,确保
AggregatedValue(CPU平均值)与阈值“小于0.6”的映射正确 - 若选择“基于结果数”,需确认触发条件设置为“当结果数大于0时触发”
- 若选择“基于度量值”,确保
3. 重复Uptime筛选导致数据丢失
查询中两次对Heartbeat执行Uptime计算和筛选,在警报动态时间窗口下,可能过滤掉符合条件的虚拟机。
- 解决方法:简化查询逻辑,仅保留一次Uptime筛选:
let timeWindow = $StartTime; let liveMachines = Heartbeat | where TimeGenerated >= timeWindow | summarize heartbeatCount = count() by Computer | extend Uptime = heartbeatCount / 1440.0 // 用浮点数避免整数除法误差 | where Uptime > 3 | project Computer, Uptime; let cpuMetrics = InsightsMetrics | where Computer in (liveMachines) | where Namespace == "Processor" and Name == "UtilizationPercentage" | where TimeGenerated between (timeWindow .. $EndTime) | extend ResourceGroup = split(_ResourceId,"/")[4] | project Computer, CPUUtilization = Val, ResourceGroup, SubscriptionID = _SubscriptionId, ResourceID = _ResourceId; cpuMetrics | summarize AvgCPU = avg(CPUUtilization) by Computer, ResourceGroup, SubscriptionID, ResourceID, Uptime = toscalar(liveMachines | where Computer == currentComputer | project Uptime) | where AvgCPU < 0.6 | project Computer, ResourceGroup, AvgCPU = round(AvgCPU,2), SubscriptionID, ResourceID, Uptime
4. 权限不足导致查询无结果
警报规则的执行账号可能缺少Log Analytics工作区读取权限,或虚拟机监控数据读取权限,导致执行时返回空结果。
- 解决方法:检查警报“运行方式”账号,确保其拥有Log Analytics工作区读取者权限,以及目标虚拟机的监控读取者权限。
5. 数据延迟导致查询未命中
InsightsMetrics或Heartbeat数据存在同步延迟,警报执行时最新数据尚未入库,导致查询无结果。
- 解决方法:在警报配置中设置适当的“延迟”(比如15分钟),确保数据完全同步后再执行查询。
内容的提问来源于stack exchange,提问作者Logan
相关产品推荐
相关产品推荐

