如何在Prometheus规则测试中正确输入container_last_seen的时间戳序列值?
如何在Prometheus规则测试中正确输入container_last_seen的时间戳序列值
问题背景
在Mac上测试Linux生产环境的Prometheus告警规则时,构造cadvisor exporter的container_last_seen指标序列触发prod_container_crashing告警时遇到格式错误:
- 输入Unix时间戳格式
'1563968832+0x61',promtool报错:parse error: unexpected number "0" in series values - 输入时长格式
'0h+1mx60',报错:parse error: unexpected duration "1h" in series values
相关配置如下:
告警规则文件(alertrules.yaml)
- name: containers interval: 15s rules: - alert: prod_container_crashing expr: | count by (instance, container_label_com_docker_swarm_service_name) ( count_over_time(container_last_seen{container_label_com_docker_swarm_service_name!="",env="prod"}[15m]) ) - 1 > 2 for: 5m labels: service: prod type: container severity: critical annotations: summary: "pdce {{ $labels.container_label_com_docker_swarm_service_name }}" description: "{{ $labels.container_label_com_docker_swarm_service_name }} in prod cluster on {{ $labels.instance }} is crashing"
测试文件(alertrules_test.yml)
rule_files: - alertrules.yml evaluation_interval: 1m tests: - name: container_tests interval: 15s input_series: - series: | container_last_seen{container_label_com_docker_swarm_service_name="service1",env="prod",instance="10.0.0.1"} values: '1563968832+0x61' alert_rule_test: - eval_time: 15m alertname: prod_container_crashing exp_alerts: - exp_labels: service: prod type: container severity: critical exp_annotations: summary: prod service1 description: service1 in prod cluster on 10.0.0.1 is crashing
解决方案
container_last_seen是gauge类型的Unix时间戳指标,在promtool测试配置中,values字段需使用纯数值格式的时间戳,增量语法要遵循数值规则:
正确写法示例
要模拟容器反复重启(每次重启时container_last_seen更新为新时间戳),可构造如下序列:
input_series: - series: | container_last_seen{container_label_com_docker_swarm_service_name="service1",env="prod",instance="10.0.0.1"} # 初始时间戳,每15s(测试interval)更新一次,生成4次变化(满足告警规则count_over_time结果-1>2的条件) values: '1690000000+1000x4'
关键说明
- 拒绝时长格式:
container_last_seen存储的是Unix时间戳数值,不是时长,promtool会直接拒绝时长格式输入。 - 数值增量语法:采用
初始值+增量x次数格式,增量为时间戳的差值(比如每次加1000代表时间推进1秒),次数对应测试周期内的样本数量。 - 匹配告警逻辑:告警规则要求15分钟内
container_last_seen至少有4次变化(count_over_time返回4,4-1=3>2),因此需要生成至少4个不同的时间戳样本。
修改后的完整测试文件
rule_files: - alertrules.yml evaluation_interval: 1m tests: - name: container_tests interval: 15s input_series: - series: | container_last_seen{container_label_com_docker_swarm_service_name="service1",env="prod",instance="10.0.0.1"} # 初始时间戳1690000000,每次加1000,生成4个样本,满足告警规则的样本数量要求 values: '1690000000+1000x4' alert_rule_test: - eval_time: 15m alertname: prod_container_crashing exp_alerts: - exp_labels: service: prod type: container severity: critical exp_annotations: summary: prod service1 description: service1 in prod cluster on 10.0.0.1 is crashing
内容的提问来源于stack exchange,提问作者Kim
相关产品推荐
相关产品推荐

