You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用ARM模板配置Azure Monitor中的新版警报?

嘿,刚好我之前研究过用ARM模板配置Azure Monitor的新版统一警报,给你梳理下具体步骤和示例,应该能帮到你😎

用ARM模板配置Azure Monitor新版警报指南

首先得明确,新版Azure Monitor警报主要依赖两个核心资源类型:动作组(Action Groups)和具体的警报规则(指标警报用Microsoft.Insights/metricAlerts,日志警报用Microsoft.Insights/logAlerts)。动作组负责定义警报触发后的通知/响应方式,警报规则则设置触发条件。

1. 先创建动作组(Action Group)

动作组是警报的“响应中枢”,你可以在这里配置邮件、短信、Webhook等通知方式。下面是一个基础的动作组ARM模板示例:

{
  "$schema": "https://schema.management.azure.com/schemas/2019-04-01/deploymentTemplate.json#",
  "contentVersion": "1.0.0.0",
  "resources": [
    {
      "type": "Microsoft.Insights/actionGroups",
      "apiVersion": "2021-09-01",
      "name": "my-production-action-group",
      "location": "global",
      "properties": {
        "groupShortName": "ProdAG",
        "enabled": true,
        "emailReceivers": [
          {
            "name": "dev-team-email",
            "emailAddress": "dev-team@yourcompany.com",
            "useCommonAlertSchema": true
          }
        ],
        "webhookReceivers": [
          {
            "name": "ops-webhook",
            "serviceUri": "https://ops.yourcompany.com/api/alerts",
            "useCommonAlertSchema": true
          }
        ]
      }
    }
  ]
}

小贴士:useCommonAlertSchema建议开启,这样所有警报的通知格式统一,后续处理起来更方便。动作组的location固定为global,因为Azure Monitor是全局服务。

2. 配置指标警报(Metric Alerts)

指标警报适用于基于Azure资源指标的监控,比如VM的CPU使用率、Web App的响应时间、SQL Server的DTU使用率等。下面是一个监控VM CPU使用率超过80%的示例:

{
  "$schema": "https://schema.management.azure.com/schemas/2019-04-01/deploymentTemplate.json#",
  "contentVersion": "1.0.0.0",
  "parameters": {
    "targetVmResourceId": {
      "type": "string",
      "description": "Resource ID of the VM you want to monitor"
    },
    "actionGroupResourceId": {
      "type": "string",
      "description": "Resource ID of your existing action group"
    }
  },
  "resources": [
    {
      "type": "Microsoft.Insights/metricAlerts",
      "apiVersion": "2021-08-01",
      "name": "vm-high-cpu-alert",
      "location": "global",
      "properties": {
        "description": "Triggers when VM CPU usage exceeds 80% for 15 minutes",
        "severity": 2, // 0=最严重, 4=最轻微
        "enabled": true,
        "scopes": [
          "[parameters('targetVmResourceId')]"
        ],
        "evaluationFrequency": "PT5M", // 每5分钟检查一次条件
        "windowSize": "PT15M", // 统计过去15分钟的指标数据
        "criteria": {
          "allOf": [
            {
              "metricName": "Percentage CPU",
              "metricNamespace": "Microsoft.Compute/virtualMachines",
              "operator": "GreaterThan",
              "threshold": 80,
              "timeAggregation": "Average",
              "dimensions": []
            }
          ],
          "odata.type": "Microsoft.Azure.Monitor.SingleResourceMultipleMetricCriteria"
        },
        "actions": [
          {
            "actionGroupId": "[parameters('actionGroupResourceId')]"
          }
        ]
      }
    }
  ]
}

关键参数说明:

  • scopes:指定要监控的资源ID,支持单个或多个资源
  • metricNamespace:对应资源的指标命名空间,比如SQL Server是Microsoft.Sql/servers,Web App是Microsoft.Web/sites
  • severity:按需设置,0级警报通常用于最紧急的故障,4级用于警告类通知

3. 配置日志警报(Log Alerts)

日志警报适用于基于Azure Monitor日志的查询,比如SQL Server的失败登录、Service Bus的消息死信、Web App的异常日志等。下面是一个监控SQL Server失败登录次数的示例:

{
  "$schema": "https://schema.management.azure.com/schemas/2019-04-01/deploymentTemplate.json#",
  "contentVersion": "1.0.0.0",
  "parameters": {
    "targetSqlServerResourceId": {
      "type": "string",
      "description": "Resource ID of your SQL Server"
    },
    "actionGroupResourceId": {
      "type": "string",
      "description": "Resource ID of your existing action group"
    }
  },
  "resources": [
    {
      "type": "Microsoft.Insights/logAlerts",
      "apiVersion": "2021-02-01-preview",
      "name": "sql-failed-logins-alert",
      "location": "global",
      "properties": {
        "description": "Triggers when there are more than 5 failed logins in 10 minutes",
        "severity": 3,
        "enabled": true,
        "scopes": [
          "[parameters('targetSqlServerResourceId')]"
        ],
        "evaluationFrequency": "PT10M",
        "windowSize": "PT10M",
        "criteria": {
          "query": "AzureDiagnostics | where ResourceProvider == 'MICROSOFT.SQL' and Category == 'SQLSecurityAuditEvents' and event_class_s == '18456' | summarize FailedLogins = count() by bin(TimeGenerated, 5m)",
          "timeAggregation": "Count",
          "operator": "GreaterThan",
          "threshold": 5,
          "odata.type": "Microsoft.Azure.Monitor.LogQueryCriteria"
        },
        "actions": [
          {
            "actionGroupId": "[parameters('actionGroupResourceId')]"
          }
        ]
      }
    }
  ]
}

关键说明:

  • query:这里是Kusto查询语句,你可以根据需要自定义查询逻辑,比如过滤特定日志类型、统计聚合结果
  • apiVersion:目前用2021-02-01-preview已经很稳定,后续有正式版可以替换

一些实用小贴士

  • 可以把动作组和警报规则放在同一个ARM模板里部署,也可以分开部署,只要确保引用的资源ID正确即可
  • 针对Service Bus、Web App等其他资源,只需要替换metricNamespace和metricName(指标警报)或者调整query(日志警报)即可
  • 新版警报支持多维度过滤,比如在指标警报的dimensions里指定按VM实例ID、Service Bus队列名称等维度监控
  • 部署前可以用az deployment group validate命令验证模板语法是否正确

内容的提问来源于stack exchange,提问作者Pradeep

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:22:42