VMSS扩缩容时资源健康告警频发致大量邮件,求优化方案
优化VMSS扩缩容导致的资源健康告警噪音方案
针对VMSS扩缩容引发的大量告警邮件问题,完全可以通过调整告警配置过滤正常操作带来的噪音,以下是具体优化方案和配置调整示例:
核心思路
VMSS扩缩容时的资源健康状态变更多为短暂的正常操作(如实例启动/删除时的临时状态波动),而非真正的资源故障。通过添加状态持续时间过滤、缩小监控范围或按资源类型拆分规则,即可有效减少无效告警。
方案一:添加状态持续时间过滤(最推荐)
通过配置告警触发的最小状态持续时间,仅当资源异常状态持续超过设定时长(如5分钟)才发送告警,过滤掉扩缩容时的瞬时状态变更。
配置调整:
- 修改
activitylog.tfvars,增加duration参数(ISO 8601格式,例如PT5M代表5分钟):
resource_health_alerts = { name = "resource-health-alert" description = "Monitors resource health (triggers only if state persists for 5 minutes)" category = "ResourceHealth" enabled = true current = ["Degraded", "Unavailable"] previous = ["Available"] reason = ["PlatformInitiated"] duration = "PT5M" # 新增:异常状态需持续5分钟才触发告警 }
- 更新
main.tf的告警规则,引入duration变量:
resource "azurerm_monitor_activity_log_alert" "this" { name = local.full_alert_name resource_group_name = var.resource_group_name description = var.alert_description scopes = local.alert_scope enabled = var.alert_enabled action { action_group_id = var.alert_action_group_id } criteria { category = var.alert_category resource_health { current = var.current previous = var.previous reason = var.reason duration = var.duration # 新增:绑定持续时间变量 } } lifecycle { ignore_changes = [ tags ] } }
- 在
variables.tf中补充duration变量定义(若未存在):
variable "duration" { type = string description = "Minimum duration the resource must stay in abnormal state to trigger alert (ISO 8601 format)" }
方案二:按资源类型拆分告警规则
将监控范围拆分,针对VMSS/VM和其他资源设置不同的告警策略:非VM类资源保持原有配置,VM/VMSS资源设置更长的持续时间阈值,或直接排除VMSS资源的监控。
示例配置(拆分规则):
- 针对非VM/VMSS资源的告警规则:
resource "azurerm_monitor_activity_log_alert" "non_vm_resource_health" { name = "${local.full_alert_name}-non-vm" resource_group_name = var.resource_group_name description = "Monitors resource health for non-VM/VMSS resources" scopes = local.alert_scope enabled = var.alert_enabled action { action_group_id = var.alert_action_group_id } criteria { category = var.alert_category resource_type = ["Microsoft.Sql/servers/databases", "Microsoft.Storage/storageAccounts"] # 指定需要监控的资源类型 resource_health { current = var.current previous = var.previous reason = var.reason } } lifecycle { ignore_changes = [ tags ] } }
- 针对VM/VMSS资源的告警规则(设置更长的持续时间):
resource "azurerm_monitor_activity_log_alert" "vm_resource_health" { name = "${local.full_alert_name}-vm" resource_group_name = var.resource_group_name description = "Monitors VM/VMSS resource health (triggers only if state persists for 10 minutes)" scopes = local.alert_scope enabled = var.alert_enabled action { action_group_id = var.alert_action_group_id } criteria { category = var.alert_category resource_type = ["Microsoft.Compute/virtualMachines", "Microsoft.Compute/virtualMachineScaleSets"] resource_health { current = var.current previous = var.previous reason = var.reason duration = "PT10M" # 针对VM/VMSS设置更长的触发阈值 } } lifecycle { ignore_changes = [ tags ] } }
方案三:缩小监控范围
如果不需要订阅级全量监控,可将告警的scope限定为特定资源组或单个资源,直接排除VMSS所在的资源组:
locals { full_alert_name = "ar-cigna-${var.alert_name}" alert_scope = [ "/subscriptions/${var.target_subscription_id}/resourceGroups/your-monitored-resource-group" # 仅包含需要监控的资源组,排除VMSS所在的RG ] }
内容的提问来源于stack exchange,提问作者samba
相关产品推荐
相关产品推荐

