使用Terraform配置基于查询表达式的CloudWatch告警
问题原因分析
你遇到的核心问题是:使用Metric Insights表达式(即SELECT语句)创建CloudWatch告警时,AWS要求必须指定指标的采集周期(period),但你当前的Terraform配置未在metric_query块中声明该参数;同时若你的Terraform AWS Provider版本过低,可能会出现无法在metric_query块中添加period的兼容性问题。
正确配置方案
1. 确保Terraform AWS Provider版本符合要求
首先确认你的provider "aws"版本不低于3.20.0(该版本开始完整支持Metric Insights告警配置),可在terraform块中指定版本:
terraform { required_providers { aws = { source = "hashicorp/aws" version = ">= 3.20.0" } } }
2. 修正CloudWatch告警资源配置
在metric_query块中添加period参数,值需与CloudWatch Agent采集磁盘指标的周期一致(默认是60秒,若你修改过采集周期则对应调整),同时注意所有数值类型参数需使用数值而非字符串类型:
resource "aws_cloudwatch_metric_alarm" "disk_usage_alarm" { alarm_name = "Disk usage alarm on MY_HOST" alarm_description = "One or more disks on MY_HOST are over 65% capacity" comparison_operator = "GreaterThanOrEqualToThreshold" threshold = 65 evaluation_periods = 2 datapoints_to_alarm = 1 treat_missing_data = "missing" actions_enabled = false insufficient_data_actions = [] alarm_actions = [] ok_actions = [] metric_query { id = "q1" label = "Maximum disk_used_percentage for all disks on Host MY_HOST" return_data = true expression = "SELECT MAX(disk_used_percent) FROM CWAgent WHERE host = 'MY_HOST'" period = 60 } }
补充说明
- 若你之前尝试用维度指定主机无结果,是因为CWAgent上报的磁盘指标包含
host和disk等多个维度,仅指定host维度会返回该主机所有磁盘的指标集合,无法直接聚合最大值,因此使用Metric Insights的SELECT表达式是更合适的方式。 period参数必须与CloudWatch Agent的指标采集周期匹配,否则会出现数据不匹配的问题,常见取值为60、300等(单位:秒)。
内容的提问来源于stack exchange,提问作者Xecu
相关产品推荐
相关产品推荐

