如何使用Terraform实现基于Log Workspace的Azure Monitor虚拟机心跳等告警
资源选型说明
azurerm_monitor_activity_log_alert仅适用于Azure平台层面的操作日志告警(比如资源创建删除、权限变更等),不支持基于Log Analytics工作区内代理上报的业务/监控日志做自定义查询告警- 你的场景需要使用
azurerm_monitor_scheduled_query_rules_alert_v2资源实现需求,KQL查询语句可以直接定义在该资源的规则条件内,无需单独创建查询资源
支持Windows/Linux双系统的告警实现示例
示例包含两个通用基础告警规则:虚拟机心跳丢失告警、虚拟机磁盘可用空间不足告警,你可以直接基于示例调整阈值、评估周期等参数。
# 前置变量定义,可根据实际情况调整 variable "log_analytics_workspace_id" { type = string description = "现有Log Analytics工作区的资源ID" } variable "action_group_id" { type = string description = "告警触发后通知使用的动作组资源ID" } # 1. 虚拟机心跳丢失告警(5分钟无心跳上报即触发) resource "azurerm_monitor_scheduled_query_rules_alert_v2" "vm_heartbeat_missing" { name = "vm-heartbeat-missing-alert" resource_group_name = "你的资源组名称" location = "你的资源区域,和Log Analytics工作区一致" evaluation_frequency = "PT5M" # 每5分钟评估一次 window_duration = "PT5M" # 每次评估过去5分钟的数据 severity = 2 # 告警级别,0最高,4最低 scopes = [var.log_analytics_workspace_id] criteria { query = <<KQL // 查询过去时间窗口内无心跳上报的虚拟机 Heartbeat | summarize LastHeartbeat = max(TimeGenerated) by Computer, OSType | where LastHeartbeat < ago(5m) | extend HostName = Computer, OS = OSType KQL time_aggregation_method = "Count" operator = "GreaterThanOrEqual" threshold = 1 failing_periods { minimum_failing_periods_to_trigger_alert = 1 number_of_evaluation_periods = 1 } } action { action_group_id = var.action_group_id } auto_mitigation_enabled = true # 告警恢复后自动发送恢复通知 } # 2. 虚拟机磁盘可用空间不足告警(可用空间低于10%触发,兼容Windows/Linux) resource "azurerm_monitor_scheduled_query_rules_alert_v2" "vm_disk_free_space_low" { name = "vm-disk-free-space-low-alert" resource_group_name = "你的资源组名称" location = "你的资源区域,和Log Analytics工作区一致" evaluation_frequency = "PT10M" window_duration = "PT10M" severity = 2 scopes = [var.log_analytics_workspace_id] criteria { query = <<KQL // 兼容Windows/Linux磁盘可用空间查询 Perf | where (ObjectName == "LogicalDisk" and CounterName == "% Free Space") // Windows 性能计数器 or (ObjectName == "Logical Disk" and CounterName == "% Free Space") // Linux 性能计数器 | where InstanceName != "_Total" and InstanceName !startswith "HarddiskVolume" // 过滤掉总容量和系统隐藏卷 | summarize AvgFreeSpace = avg(CounterValue) by Computer, OSType, InstanceName | where AvgFreeSpace < 10 // 可用空间低于10%,可自行调整阈值 | extend HostName = Computer, OS = OSType, DiskMountPoint = InstanceName, FreeSpacePercent = AvgFreeSpace KQL time_aggregation_method = "Count" operator = "GreaterThanOrEqual" threshold = 1 failing_periods { minimum_failing_periods_to_trigger_alert = 2 number_of_evaluation_periods = 3 } } action { action_group_id = var.action_group_id } auto_mitigation_enabled = true }
参数调整说明
- 可以修改
evaluation_frequency和window_duration调整告警的评估频率和统计周期,格式为ISO 8601时长格式,PT1H代表1小时,PT1M代表1分钟 - KQL语句中的阈值可以按需调整,比如磁盘可用空间阈值可以改为5%或者20%
failing_periods块可以调整触发告警的连续失败次数,避免偶发数据抖动导致误告警
内容的提问来源于stack exchange,提问作者SNielsen
相关产品推荐
相关产品推荐

