You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Alert-v1低CPU利用率警报不触发问题求助

Azure Alert-v1警报无法触发排查求助

我使用以下自定义Kusto查询筛选运行时长超过3天且CPU利用率较低的虚拟机,查询单独执行时可正常返回结果,但基于此创建Azure Alert-v1警报后,警报始终无法触发。附上查询语句及警报配置截图说明,恳请协助排查问题。

let startdate = ago(7d);
let enddate = now();
let livemachine = Heartbeat 
| where TimeGenerated > ago(7d) 
| summarize AggregatedValue = count() by Computer 
| extend Uptime = AggregatedValue / 1440
| where Uptime > 3 
| project Computer ;
let cpu = InsightsMetrics
    | where Computer in (livemachine)
    | where Namespace contains "Processor" and Name contains "UtilizationPercentage"
    | where TimeGenerated between (startdate .. enddate)
    | project
        startdate,
        enddate,
        TimeGenerated,
        Computer,
        CPUUtilization = Val,
        SubscriptionID = _SubscriptionId,
        ResourceID = _ResourceId
| extend ResourceGroup = split(ResourceID,"/")[4];
cpu | summarize AggregatedValue = avg(CPUUtilization)
    by
    bin(TimeGenerated, 7d),
    Computer,
    tostring(ResourceGroup),
    SubscriptionID,
    ResourceID,
    startdate,
    enddate
| join kind=inner (Heartbeat | where TimeGenerated > ago(7d)
| summarize AggregatedValue = count() by Computer 
| extend Uptime = AggregatedValue / 1440
| where Uptime > 3 
| project Computer, Uptime ) on Computer
| where AggregatedValue < 0.6
| project
    startdate,
    enddate,
    Computer,
    ResourceGroup,
    round(AggregatedValue, 2),
    SubscriptionID,
    ResourceID,
    TimeGenerated,
    Uptime
| summarize arg_max(TimeGenerated, *) by Computer

警报配置截图说明

  • 截图1:展示警报规则的基础配置,包含名称、描述、资源范围等内容
  • 截图2:展示信号逻辑配置,包含查询关联、周期、频率、触发条件等参数
  • 截图3:展示操作组相关配置内容

排查方向及解决方法

1. 硬编码时间范围与警报动态窗口冲突

查询中硬编码了startdate = ago(7d)和enddate = now(),但Azure警报会自动注入$StartTime和$EndTime参数,对应警报配置的周期和频率时间窗口。硬编码的时间会覆盖动态窗口,导致警报执行时的查询范围与手动执行不一致。

  • 解决方法:替换为警报内置参数:
    let startdate = $StartTime;
    let enddate = $EndTime;
    
    同时确保警报配置的周期设置为7天,频率按需调整(比如每天执行一次)。

2. 聚合逻辑与警报触发条件不匹配

查询最后用arg_max(TimeGenerated, *) by Computer仅保留每个虚拟机的最新记录,若警报触发条件依赖特定字段的聚合逻辑,可能无法识别有效触发项。

  • 解决方法:检查警报触发条件配置:
    • 若选择“基于度量值”,确保AggregatedValue(CPU平均值)与阈值“小于0.6”的映射正确
    • 若选择“基于结果数”,需确认触发条件设置为“当结果数大于0时触发”

3. 重复Uptime筛选导致数据丢失

查询中两次对Heartbeat执行Uptime计算和筛选,在警报动态时间窗口下,可能过滤掉符合条件的虚拟机。

  • 解决方法:简化查询逻辑,仅保留一次Uptime筛选:
    let timeWindow = $StartTime;
    let liveMachines = Heartbeat 
    | where TimeGenerated >= timeWindow
    | summarize heartbeatCount = count() by Computer 
    | extend Uptime = heartbeatCount / 1440.0  // 用浮点数避免整数除法误差
    | where Uptime > 3 
    | project Computer, Uptime;
    let cpuMetrics = InsightsMetrics
        | where Computer in (liveMachines)
        | where Namespace == "Processor" and Name == "UtilizationPercentage"
        | where TimeGenerated between (timeWindow .. $EndTime)
        | extend ResourceGroup = split(_ResourceId,"/")[4]
        | project Computer, CPUUtilization = Val, ResourceGroup, SubscriptionID = _SubscriptionId, ResourceID = _ResourceId;
    cpuMetrics
    | summarize AvgCPU = avg(CPUUtilization) by Computer, ResourceGroup, SubscriptionID, ResourceID, Uptime = toscalar(liveMachines | where Computer == currentComputer | project Uptime)
    | where AvgCPU < 0.6
    | project Computer, ResourceGroup, AvgCPU = round(AvgCPU,2), SubscriptionID, ResourceID, Uptime
    

4. 权限不足导致查询无结果

警报规则的执行账号可能缺少Log Analytics工作区读取权限,或虚拟机监控数据读取权限,导致执行时返回空结果。

  • 解决方法:检查警报“运行方式”账号,确保其拥有Log Analytics工作区读取者权限,以及目标虚拟机的监控读取者权限。

5. 数据延迟导致查询未命中

InsightsMetrics或Heartbeat数据存在同步延迟,警报执行时最新数据尚未入库,导致查询无结果。

  • 解决方法:在警报配置中设置适当的“延迟”(比如15分钟),确保数据完全同步后再执行查询。

内容的提问来源于stack exchange,提问作者Logan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 08:04:52