寻求Scheduled Query Alert实用查询示例及Bicep配置指导
一、Application Insights 场景查询示例
针对应用性能、异常、依赖调用等核心场景,以下是实用的KQL查询及告警配置建议:
异常率过高告警
用途:当未处理异常占总请求比例超过阈值时触发
查询语句:requests | where timestamp > ago(5m) | extend hasException = tobool(customDimensions['hasException'] or resultCode startswith '5') | summarize totalRequests = count(), exceptionRequests = countif(hasException) by bin(timestamp, 1m) | extend exceptionRate = exceptionRequests * 100.0 / totalRequests | where exceptionRate > 5 // 异常率超过5%配置参考:查询周期5分钟,频率5分钟,阈值设为>0,严重等级标记为警告。
请求失败率过高告警
用途:HTTP 5xx/4xx错误占比超出业务容忍度时触发
查询语句:requests | where timestamp > ago(5m) | extend isFailed = resultCode startswith '4' or resultCode startswith '5' | summarize total = count(), failed = countif(isFailed) by bin(timestamp, 1m) | extend failureRate = failed * 100.0 / total | where failureRate > 10 // 失败率超过10%配置参考:查询周期5分钟,频率5分钟,阈值>0,严重等级根据业务设定为警告/错误。
依赖调用超时告警
用途:外部依赖(如数据库、API)响应超时请求占比过高时触发
查询语句:dependencies | where timestamp > ago(5m) | extend isTimeout = duration > 3000 // 3秒视为超时,可根据业务调整 | summarize totalCalls = count(), timeoutCalls = countif(isTimeout) by bin(timestamp, 1m) | extend timeoutRate = timeoutCalls * 100.0 / totalCalls | where timeoutRate > 8 // 超时率超过8%配置参考:查询周期5分钟,频率5分钟,阈值>0,严重等级为警告。
错误日志激增告警
用途:短时间内错误日志数量较基线翻倍时触发
查询语句:traces | where timestamp > ago(10m) | where severityLevel == 3 // 筛选Error级日志 | summarize errorCount = count() by bin(timestamp, 1m) | join kind=inner ( traces | where timestamp > ago(20m) and timestamp < ago(10m) | where severityLevel == 3 | summarize baselineCount = count() by bin(timestamp, 1m) | summarize avgBaseline = avg(baselineCount) ) on $left.empty = $right.empty | where errorCount > avgBaseline * 2 // 错误量超过基线2倍配置参考:查询周期10分钟,频率5分钟,阈值>0,严重等级为错误。
二、Network Watcher 场景查询示例
针对网络连接、流量管控、线路健康等场景:
VPN连接中断告警
用途:VPN网关连接状态变为断开时触发
查询语句:AzureDiagnostics | where ResourceType == 'VIRTUALNETWORKGATEWAYS' | where Category == 'GatewayDiagnosticLog' | where OperationName == 'TunnelHealth' | where status_s == 'Disconnected' | where timestamp > ago(5m) | summarize count() by ResourceName, bin(timestamp, 1m)配置参考:查询周期5分钟,频率5分钟,阈值>0,严重等级为错误。
NSG拒绝流量激增告警
用途:NSG短时间内拒绝请求量突增时触发
查询语句:AzureNetworkAnalytics_CL | where SubType_s == 'FlowLog' | where Action_s == 'Deny' | where timestamp > ago(5m) | summarize denyCount = count() by NSGName_s, bin(timestamp, 1m) | where denyCount > 100 // 5分钟内拒绝超过100次,可调整配置参考:查询周期5分钟,频率5分钟,阈值>0,严重等级为警告。
ExpressRoute线路健康异常告警
用途:ExpressRoute线路状态变为Down时触发
查询语句:AzureDiagnostics | where ResourceType == 'EXPRESSROUTECIRCUITS' | where Category == 'PeeringDiagnosticLog' | where status_s == 'Down' | where timestamp > ago(10m) | summarize count() by ResourceName, bin(timestamp, 1m)配置参考:查询周期10分钟,频率5分钟,阈值>0,严重等级为错误。
三、Azure Sentinel 场景查询示例
针对安全事件、异常访问、威胁检测等场景:
异地异常登录告警
用途:同一账号短时间内从多个地理位置登录时触发
查询语句:SigninLogs | where timestamp > ago(1h) | where ResultType == 0 // 筛选成功登录事件 | summarize locations = make_set(Location) by UserPrincipalName, bin(timestamp, 15m) | where array_length(locations) > 1 // 15分钟内来自多个地区配置参考:查询周期1小时,频率15分钟,阈值>0,严重等级为高。
高/中等级恶意软件检测告警
用途:终端检测到高/中风险恶意软件时触发
查询语句:DeviceEvents | where timestamp > ago(30m) | where ActionType == 'AntivirusDetection' | where ThreatSeverity in ('High', 'Medium') | summarize count() by DeviceName, ThreatName, bin(timestamp, 10m)配置参考:查询周期30分钟,频率10分钟,阈值>0,严重等级为高。
非授权访问敏感存储告警
用途:指定授权外的用户访问敏感存储账户时触发
查询语句:StorageBlobLogs | where timestamp > ago(1h) | where OperationName == 'GetBlob' | where AccountName contains 'sensitive' // 替换为你的敏感存储关键词 | where UserPrincipalName !in ('authorized.user@domain.com', 'admin@domain.com') // 替换为授权用户列表 | summarize count() by UserPrincipalName, bin(timestamp, 15m)配置参考:查询周期1小时,频率15分钟,阈值>0,严重等级为高。
四、Bicep 模板配置参考
以下是填充查询后的完整Scheduled Query Rule Bicep示例,关键参数可按需替换:
resource scheduledQueryRule 'Microsoft.Insights/scheduledQueryRules@2022-06-01' = { name: 'app-insights-exception-rate-alert' location: resourceGroup().location properties: { displayName: 'Application Insights 异常率过高告警' description: '当应用异常率超过5%时触发告警' enabled: true source: { query: ''' requests | where timestamp > ago(5m) | extend hasException = tobool(customDimensions['hasException'] or resultCode startswith '5') | summarize totalRequests = count(), exceptionRequests = countif(hasException) by bin(timestamp, 1m) | extend exceptionRate = exceptionRequests * 100.0 / totalRequests | where exceptionRate > 5 ''' dataSourceId: resourceId('Microsoft.Insights/components', 'your-app-insights-name') // 替换为你的Application Insights资源ID queryType: 'ResultCount' } schedule: { frequencyInMinutes: 5 timeWindowInMinutes: 5 } action: { aznsAction: { actionGroup: [resourceId('Microsoft.Insights/actionGroups', 'your-action-group-name')] // 替换为你的动作组ID emailSubject: '应用异常率过高告警' } severity: 2 // 严重等级:0-严重,1-错误,2-警告,3-信息,4-详细 throttlingInMinutes: 5 // 告警抑制时间,避免重复触发 } criteria: { operator: 'GreaterThan' threshold: 0 timeAggregation: 'Count' } } }
关键参数说明:
dataSourceId: 目标监控资源的ID(Application Insights/Log Analytics工作区等)frequencyInMinutes: 告警查询的执行周期timeWindowInMinutes: 查询覆盖的历史时间范围severity: 根据业务影响程度设置告警等级throttlingInMinutes: 抑制重复告警的时间窗口,避免短时间内多次推送相同告警
内容的提问来源于stack exchange,提问作者Jennifer M.

