Terraform创建CloudWatch磁盘告警持续insufficient_data如何解决
问题概述
- 多台运行遗留业务的Windows EC2实例,每台挂载200GB D盘辅助卷存储应用与日志
- 通过Terraform创建CloudWatch磁盘空间告警,资源创建成功后所有告警持续处于
insufficient_data状态无法正常触发 - 已尝试调整dimensions参数,添加
path = /、device = xvda配置项,问题未解决 - 原有Terraform配置片段如下:
data "aws_instances" "this" { filter { name = "image-id" values = [data.aws_ami.this["windows"].image_id] } } resource "aws_cloudwatch_metric_alarm" "this" { for_each = toset(data.aws_instances.this.ids) alarm_name = "Disk-space-${each.value}" comparison_operator = "LessThanOrEqualToThreshold" evaluation_periods = "1" metric_name = "LogicalDisk % Free Space" namespace = "CWAgent" period = "180" statistic = "Average" threshold = "20" alarm_description = "This metric monitors free space on application drive" actions_enabled = "true" alarm_actions = ["arn:aws:sns:xxxxxxx] insufficient_data_actions = [] #treat_missing_data = "notBreaching" dimensions = { InstanceId = each.value Instance = "D:" } }
排查方向
- 确认指标来源是否正常:配置中使用的
CWAgent命名空间不属于EC2默认上报的指标集,必须在目标实例上安装、正确配置CloudWatch Agent后才会上报对应指标。EC2默认自带的磁盘指标归属AWS/EC2命名空间,仅上报根卷状态,无法采集单独挂载的D盘逻辑磁盘数据。 - 核对维度配置匹配性:Windows系统下
LogicalDisk % Free Space指标的逻辑盘维度名为LogicalDisk而非Instance,之前添加的path = /、device = xvda是Linux系统磁盘指标的专属维度,在Windows环境下无效。告警配置的维度必须和Agent实际上报的维度名、取值完全一致,大小写敏感。 - 检查采集周期匹配:当前告警配置的聚合周期为180秒,需确认CloudWatch Agent配置的磁盘指标采集间隔小于等于180秒,且上报时未携带告警配置中未声明的额外自定义维度。
- 验证实例权限:目标EC2绑定的IAM实例角色必须附加
CloudWatchAgentServerPolicy托管策略,否则Agent无权限向CloudWatch上报指标数据。 - 修正配置语法错误:原配置中
alarm_actions字段的SNS ARN缺失闭合双引号,属于语法瑕疵。
解决方案
- 所有目标Windows实例安装CloudWatch Agent,在Agent配置中开启D盘逻辑磁盘指标采集,参考配置段如下:
{ "metrics": { "metrics_collected": { "LogicalDisk": { "measurement": ["% Free Space"], "resources": ["D:"], "append_dimensions": { "InstanceId": "${aws:InstanceId}" } } } } }
配置完成后重启CloudWatch Agent服务,确保服务处于运行状态。
- 修正Terraform告警配置,调整错误的维度项,补全语法缺失部分,修正后配置如下:
resource "aws_cloudwatch_metric_alarm" "this" { for_each = toset(data.aws_instances.this.ids) alarm_name = "Disk-space-${each.value}" comparison_operator = "LessThanOrEqualToThreshold" evaluation_periods = "1" metric_name = "LogicalDisk % Free Space" namespace = "CWAgent" period = "180" statistic = "Average" threshold = "20" alarm_description = "This metric monitors free space on application drive" actions_enabled = "true" alarm_actions = ["arn:aws:sns:xxxxxxx"] insufficient_data_actions = [] dimensions = { InstanceId = each.value LogicalDisk = "D:" } }
执行terraform apply更新告警配置。
验证步骤
- 进入CloudWatch控制台指标浏览页,选择
CWAgent命名空间,按目标实例ID筛选,确认存在LogicalDisk % Free Space指标,且LogicalDisk = D:维度下有连续数据点上报 - 等待2-3个采集周期(约5-10分钟),告警状态会自动从
insufficient_data切换为OK或ALARM正常状态。
内容的提问来源于stack exchange,提问作者CMR H
相关产品推荐
相关产品推荐

