使用Terraform自动化配置GCP Stackdriver告警与通知的问题求助
解决Terraform配置Stackdriver告警与通知关联的问题
你已经搞定了通知渠道和正常运行时间检查的基础配置,接下来只需要添加告警策略资源就能把两者关联起来,实现当uptime检查失败时自动触发通知。下面是直接结合你现有代码的完整配置方案:
完整配置代码
# 你已有的通知渠道配置 resource "google_monitoring_notification_channel" "basic" { display_name = "Test Notification Channel" type = "email" labels = { email_address = "fakeid007@gmail.com" } project = "department1" } # 你已有的uptime检查配置 resource "google_monitoring_uptime_check_config" "http" { display_name = "01 - Website uptime check [global]" timeout = "10s" period = "60s" project = "department1" http_check { path = "/" port = "80" mask_headers = null use_ssl = null validate_ssl = null request_method = "GET" } monitored_resource { type = "uptime_url" labels = { host = "35.184.98.16" } } } # 新增告警策略:关联uptime检查与通知渠道 resource "google_monitoring_alert_policy" "uptime_failure_alert" { display_name = "Uptime Check Failure Alert" project = "department1" # 定义告警触发条件:当uptime检查失败时 conditions { display_name = "Uptime Check Failed" condition_threshold { # 精准匹配你创建的uptime检查,用Terraform插值避免手动输入错误 filter = "metric.type=\"monitoring.googleapis.com/uptime_check/check_failed\" AND resource.type=\"uptime_url\" AND resource.label.\"host\"=\"${google_monitoring_uptime_check_config.http.monitored_resource.0.labels.host}\"" comparison = "COMPARISON_GT" threshold_value = 0 duration = "0s" # 检查失败立即触发告警 # 配置触发次数:这里设置1次失败就告警,可按需调整 trigger { count = 1 } } } # 绑定你之前创建的邮件通知渠道 notification_channels = [google_monitoring_notification_channel.basic.id] # 告警严重级别,可选值:CRITICAL、ERROR、WARNING、INFO、DEBUG severity = "CRITICAL" }
关键细节说明
- 条件过滤:通过
filter规则精准关联你创建的特定uptime检查,用Terraform插值语法直接引用现有资源的host参数,避免手动输入错误。 - 触发逻辑:
check_failed指标在检查正常时为0,失败时为1,所以设置threshold_value = 0+COMPARISON_GT就能在失败时立即触发告警。 - 通知绑定:
notification_channels直接引用你已创建的通知渠道ID,确保告警触发时自动发送邮件。
验证步骤
运行terraform apply完成部署后,你可以在Stackdriver控制台确认:
- 「告警」>「告警策略」列表中会出现你配置的
Uptime Check Failure Alert - 手动模拟站点故障(比如关闭目标服务),测试是否能收到告警邮件通知
内容的提问来源于stack exchange,提问作者d.s
相关产品推荐
相关产品推荐

