You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Terraform创建CloudWatch磁盘告警持续insufficient_data如何解决

问题概述
  • 多台运行遗留业务的Windows EC2实例,每台挂载200GB D盘辅助卷存储应用与日志
  • 通过Terraform创建CloudWatch磁盘空间告警,资源创建成功后所有告警持续处于insufficient_data状态无法正常触发
  • 已尝试调整dimensions参数,添加path = /、device = xvda配置项,问题未解决
  • 原有Terraform配置片段如下:
data "aws_instances" "this" {
  filter {
    name   = "image-id"
    values = [data.aws_ami.this["windows"].image_id]
  }
}

resource "aws_cloudwatch_metric_alarm" "this" {
  for_each                  = toset(data.aws_instances.this.ids)
  alarm_name                = "Disk-space-${each.value}"
  comparison_operator       = "LessThanOrEqualToThreshold"
  evaluation_periods        = "1"
  metric_name               = "LogicalDisk % Free Space"
  namespace                 = "CWAgent"
  period                    = "180"
  statistic                 = "Average"
  threshold                 = "20"
  alarm_description         = "This metric monitors free space on application drive"
  actions_enabled           = "true"
  alarm_actions             = ["arn:aws:sns:xxxxxxx]
  insufficient_data_actions = []
  #treat_missing_data = "notBreaching"

  dimensions = {
    InstanceId = each.value
    Instance   = "D:"

  }
}
排查方向
  • 确认指标来源是否正常:配置中使用的CWAgent命名空间不属于EC2默认上报的指标集,必须在目标实例上安装、正确配置CloudWatch Agent后才会上报对应指标。EC2默认自带的磁盘指标归属AWS/EC2命名空间,仅上报根卷状态,无法采集单独挂载的D盘逻辑磁盘数据。
  • 核对维度配置匹配性:Windows系统下LogicalDisk % Free Space指标的逻辑盘维度名为LogicalDisk而非Instance,之前添加的path = /、device = xvda是Linux系统磁盘指标的专属维度,在Windows环境下无效。告警配置的维度必须和Agent实际上报的维度名、取值完全一致,大小写敏感。
  • 检查采集周期匹配:当前告警配置的聚合周期为180秒,需确认CloudWatch Agent配置的磁盘指标采集间隔小于等于180秒,且上报时未携带告警配置中未声明的额外自定义维度。
  • 验证实例权限:目标EC2绑定的IAM实例角色必须附加CloudWatchAgentServerPolicy托管策略,否则Agent无权限向CloudWatch上报指标数据。
  • 修正配置语法错误:原配置中alarm_actions字段的SNS ARN缺失闭合双引号,属于语法瑕疵。
解决方案
  1. 所有目标Windows实例安装CloudWatch Agent,在Agent配置中开启D盘逻辑磁盘指标采集,参考配置段如下:
{
  "metrics": {
    "metrics_collected": {
      "LogicalDisk": {
        "measurement": ["% Free Space"],
        "resources": ["D:"],
        "append_dimensions": {
          "InstanceId": "${aws:InstanceId}"
        }
      }
    }
  }
}

配置完成后重启CloudWatch Agent服务,确保服务处于运行状态。

  1. 修正Terraform告警配置,调整错误的维度项,补全语法缺失部分,修正后配置如下:
resource "aws_cloudwatch_metric_alarm" "this" {
  for_each                  = toset(data.aws_instances.this.ids)
  alarm_name                = "Disk-space-${each.value}"
  comparison_operator       = "LessThanOrEqualToThreshold"
  evaluation_periods        = "1"
  metric_name               = "LogicalDisk % Free Space"
  namespace                 = "CWAgent"
  period                    = "180"
  statistic                 = "Average"
  threshold                 = "20"
  alarm_description         = "This metric monitors free space on application drive"
  actions_enabled           = "true"
  alarm_actions             = ["arn:aws:sns:xxxxxxx"]
  insufficient_data_actions = []

  dimensions = {
    InstanceId  = each.value
    LogicalDisk = "D:"
  }
}

执行terraform apply更新告警配置。

验证步骤
  • 进入CloudWatch控制台指标浏览页,选择CWAgent命名空间,按目标实例ID筛选,确认存在LogicalDisk % Free Space指标,且LogicalDisk = D:维度下有连续数据点上报
  • 等待2-3个采集周期(约5-10分钟),告警状态会自动从insufficient_data切换为OK或ALARM正常状态。

内容的提问来源于stack exchange,提问作者CMR H

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 12:33:06