如何将K8s Pod的custom.io注解自动注入Prometheus告警通知?
实现方案说明
方案1:调整kube-state-metrics配置(推荐,开发量最小)
你之前配置的relabel_configs没有生效,核心原因是该配置作用于Prometheus直接采集Pod业务指标的任务,而你告警使用的kube_pod_container_status_restarts_total指标是kube-state-metrics暴露的,元数据配置需要在kube-state-metrics侧调整:
- 给kube-state-metrics添加启动参数,允许暴露自定义前缀的Pod注解作为指标标签:
--metric-annotations-allowlist=pods=custom.io/*
配置后所有custom.io开头的Pod注解会自动转为kube_pod系列指标的标签,标签格式为annotation_custom_io_<注解名>,比如custom.io/runbook对应的标签名为annotation_custom_io_runbook。
2. 调整告警规则直接引用标签即可:
- alert: RestartsCountWarning annotations: description: Warning - Pod has restarted at least 3 times summary: Pod restarting (>3) runbook: {{ $labels.annotation_custom_io_runbook }} logs: {{ $labels.annotation_custom_io_logs }} monitoring: {{ $labels.annotation_custom_io_monitoring }} owner: {{ $labels.annotation_custom_io_owner }} expr: kube_pod_container_status_restarts_total{job="kube-state-metrics",namespace=".*"} > 3 for: 5m labels: severity: warning
- 后续Slack告警模板无需调整,直接通过
{{ .Annotations.runbook }}这类语法调用即可。
说明:你之前担心的将注解作为指标标签不符合规范的问题,仅在标签为高基数的情况下存在,这类低基数的应用元数据标签属于Prometheus推荐的用法,不会造成存储或查询性能问题。
方案2:Alertmanager webhook二次加工(无指标标签侵入)
如果你完全不想在Prometheus指标中存储注解信息,可以在告警链路增加一层自定义webhook服务:
- Alertmanager触发告警后,先将告警推送到自定义webhook
- webhook根据告警中携带的
namespace、pod标签,调用K8s API查询对应Pod的custom.io前缀注解 - 把查询到的注解补充到告警的annotations字段后,再转发到Slack
该方案无需修改Prometheus侧配置,缺点是需要额外维护webhook服务,告警链路复杂度更高。
内容的提问来源于stack exchange,提问作者chepeftw
相关产品推荐
相关产品推荐

