You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure应用网关可用性SLI告警问题:月度SLI计算与查询范围限制

问题

我要实现Azure告警,当应用网关的可用性SLI低于99.9%阈值时触发。可用性SLI按当月至今数据计算,公式为:100 - (5xx状态码请求数 / 总请求数)。

我写了一段Kusto查询,按5分钟间隔计算当月滚动平均值(假设当月剩余时间可用性为100%,以此体现错误预算),代码如下:

let resolution = 5m;
let monthStart = startofmonth(datetime(now));
let monthEnd = endofmonth(datetime(now));
let now = datetime(now);
AzureDiagnostics
| where ResourceType == "APPLICATIONGATEWAYS"
    and OperationName == "ApplicationGatewayAccess"
    and TimeGenerated >= monthStart
    and TimeGenerated <= monthEnd
| summarize
    TotalRequests = count(),
    ErrorRequests = countif(httpStatus_d > 499)
    by bin(TimeGenerated, resolution)
| sort by TimeGenerated asc
| serialize Period = row_number()
| extend periodsLeft = round((monthEnd - TimeGenerated) / resolution)
| extend periodsTotal = Period + periodsLeft
| extend AvailabilityRateInPeriod = 100 - (todouble(ErrorRequests) / TotalRequests * 100)
| serialize RunningPeriodSum = row_cumsum(AvailabilityRateInPeriod)
| extend AvailabilityRateRunning = (RunningPeriodSum + (100 * periodsLeft)) / periodsTotal
| project TimeGenerated, AvailabilityRateRunning

这段查询单独使用时能正常生成数据或图表,但配置成告警时,发现告警有最大2天的回溯期,导致查询只能获取最近2天的数据,没法计算完整的当月至今SLI。想请教:

  1. 能否通过存储滚动中间值来实现告警?
  2. 有没有其他更优方案?
  3. 有没有更简洁高效的KQL写法?
解决方案

方案1:日志警报 + 累积数据存储

利用Azure Monitor的已保存搜索定期聚合全月数据,存储到自定义日志表中,告警查询时合并历史聚合数据与实时增量数据,规避回溯期限制:

  1. 创建已保存搜索,每天运行一次全月聚合查询,将结果写入自定义日志表(如MonthlySLIAggregates):
    let monthStart = startofmonth(now());
    let now = now();
    AzureDiagnostics
    | where ResourceType == "APPLICATIONGATEWAYS"
        and OperationName == "ApplicationGatewayAccess"
        and TimeGenerated >= monthStart
        and TimeGenerated <= now
    | summarize
        TotalRequests = sum(count()),
        ErrorRequests = sum(countif(httpStatus_d > 499))
        by ResourceId, bin(monthStart, 1d)
    | extend AvailabilityRate = 100 - (todouble(ErrorRequests) / TotalRequests * 100)
    | extend Month = monthStart
    | project ResourceId, Month, TotalRequests, ErrorRequests, AvailabilityRate
    | into MonthlySLIAggregates
    
  2. 配置告警时,查询自定义表获取历史数据,结合最近2天的实时数据补全计算:
    let monthStart = startofmonth(now());
    let now = now();
    // 取当月已聚合的历史数据
    let historical = MonthlySLIAggregates
    | where Month == monthStart
    | summarize Total_Hist = sum(TotalRequests), Error_Hist = sum(ErrorRequests);
    // 取最近2天的实时增量
    let realtime = AzureDiagnostics
    | where ResourceType == "APPLICATIONGATEWAYS"
        and OperationName == "ApplicationGatewayAccess"
        and TimeGenerated >= ago(2d)
        and TimeGenerated <= now
    | summarize Total_Real = count(), Error_Real = countif(httpStatus_d > 499);
    // 合并计算当月SLI
    historical
    | join kind=fullouter realtime on $left.ResourceId == $right.ResourceId
    | extend Total = coalesce(Total_Hist, 0) + coalesce(Total_Real, 0)
    | extend Error = coalesce(Error_Hist, 0) + coalesce(Error_Real, 0)
    | extend AvailabilityRate = 100 - (todouble(Error) / Total * 100)
    | project AvailabilityRate
    

方案2:改用Azure Monitor指标规则(推荐)

如果能将应用网关的请求数、错误数导出为自定义指标,可直接利用指标平台实现跨月聚合:

  1. 通过诊断设置,将AzureDiagnostics中的请求数、5xx错误数转换为自定义指标,导出到Azure Monitor指标平台。
  2. 创建指标警报规则,选择“当月至今”的时间范围,用Sum聚合总请求数与错误数,配置自定义计算规则100 - (ErrorRequests / TotalRequests * 100)作为SLI,设置阈值触发告警。
  3. 优势:指标平台支持长周期聚合,不受日志告警的回溯期限制,性能更优。

优化后的KQL写法

针对原查询简化逻辑、提升可读性:

let monthStart = startofmonth(now());
let monthEnd = endofmonth(now());
let resolution = 5m;
AzureDiagnostics
| where ResourceType == "APPLICATIONGATEWAYS"
    and OperationName == "ApplicationGatewayAccess"
    and TimeGenerated between (monthStart .. monthEnd)
| summarize
    Total = count(),
    Errors = countif(httpStatus_d > 499)
    by bin(TimeGenerated, resolution)
| order by TimeGenerated asc
| serialize
    CumulativeTotal = row_cumsum(Total),
    CumulativeErrors = row_cumsum(Errors)
| extend
    remainingPeriods = round((monthEnd - TimeGenerated) / resolution),
    cumulativeAvailability = 100 - (todouble(CumulativeErrors) / CumulativeTotal * 100)
| extend
    rollingAvailability = (cumulativeAvailability * row_number() + 100 * remainingPeriods) / (row_number() + remainingPeriods)
| project TimeGenerated, rollingAvailability

优化点:

  • 直接累计请求/错误数,避免单独计算周期分数再求和,逻辑更直观。
  • 用between简化时间范围判断。
  • 精简冗余变量,代码更紧凑。

内容的提问来源于stack exchange,提问作者devguydavid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 05:33:32