Promtool单元测试未达预期:内存使用率告警测试失败求助
内存使用率告警规则测试失败排查
告警规则定义
- alert: MemoryUsage90Percent expr: sum(container_memory_usage_bytes{job="cadvisor",container_label_io_rancher_stack_name!="", image=~"rancher.*"}) by (container_label_io_rancher_stack_name, instance, dns_name) /1024/1024/1024 *100 > 90 for: 5m labels: severity: Critical alert_owner: alerts-infra annotations: description: "Service memory has been at over 90% for 1 minute. Service: {{ $labels.container_label_io_rancher_stack_name }}, DnsName: {{ $labels.dns_name }}."
测试用例代码
# Infra Alert Tests - input_series: - series: 'container_memory_usage_bytes{job="cadvisor",container_label_io_rancher_stack_name="network-service",image="rancher.test"}' values: '0 100000000 0 100000000 100000000' alert_rule_test: - alertname: MemoryUsage90Percent # eval_time: 5m exp_alerts: - exp_labels: alertname: MemoryUsage90Percent severity: Critical alert_owner: alerts-infra exp_annotations: description: "Service memory has been at over 90% for 1 minute. Service: (), DnsName: ()."
测试失败结果
Unit Testing: tests/infra_db_alerts-tests.yml FAILED: alertname: ServiceMemoryUsage90Percent, time: 0s, exp:[ 0: Labels:{alert_owner="alerts-infra", alertname="ServiceMemoryUsage90Percent", severity="Critical"} Annotations:{description="Service memory has been at over 90% for 1 minute. Service: (), DnsName: ()."} ], got:[]
问题根源及修复方案
1. 告警规则表达式逻辑错误
当前表达式仅将内存使用字节数转换为GB后乘以100,判断是否大于90——这实际是判断使用量是否超过0.9GB,而非内存使用率(使用量/总内存)超过90%,逻辑完全错误。同时这也解释了“输入0时告警仍触发”的怪异现象(大概率是测试环境中存在其他符合标签的高内存数据,或表达式书写时的语法疏漏)。
修正后的表达式需计算使用率:
expr: sum(container_memory_usage_bytes{job="cadvisor",container_label_io_rancher_stack_name!="", image=~"rancher.*"}) by (container_label_io_rancher_stack_name, instance, dns_name) / sum(container_memory_limit_bytes{job="cadvisor",container_label_io_rancher_stack_name!="", image=~"rancher.*"}) by (container_label_io_rancher_stack_name, instance, dns_name) * 100 > 90
注:需确保
container_memory_limit_bytes指标存在,且与container_memory_usage_bytes的标签维度完全匹配,否则会出现相除失败的情况。
2. 测试用例数据未达告警阈值
当前测试输入的100000000字节仅约0.093GB,转换后为9.3,远小于90的触发条件,因此不会生成告警,导致got:[]。
需调整输入数据,结合修正后的表达式,设置使用量占总内存的90%以上:
input_series: # 总内存设为1GB(1073741824字节) - series: 'container_memory_limit_bytes{job="cadvisor",container_label_io_rancher_stack_name="network-service",image="rancher.test",instance="test-instance",dns_name="test-dns"}' values: '1073741824+0x' # 使用量设为950MB(996147200字节),使用率约90.3% - series: 'container_memory_usage_bytes{job="cadvisor",container_label_io_rancher_stack_name="network-service",image="rancher.test",instance="test-instance",dns_name="test-dns"}' values: '0 996147200 996147200 996147200 996147200'
3. 测试用例缺少必要配置
- 启用
eval_time:告警规则设置了for:5m,需等待5分钟告警才会从pending转为firing状态,测试时必须添加eval_time: 5m。 - 补充缺失标签:告警规则按
instance、dns_name分组,但测试输入的序列缺少这两个标签,导致sum后无法匹配规则维度,需在输入series中添加这两个标签。 - 修正预期注释内容:预期的
description中,Service应为network-service,DnsName应为输入的test-dns,而非空括号。
修正后的完整测试用例
# Infra Alert Tests - input_series: - series: 'container_memory_limit_bytes{job="cadvisor",container_label_io_rancher_stack_name="network-service",image="rancher.test",instance="test-instance",dns_name="test-dns"}' values: '1073741824+0x' - series: 'container_memory_usage_bytes{job="cadvisor",container_label_io_rancher_stack_name="network-service",image="rancher.test",instance="test-instance",dns_name="test-dns"}' values: '0 996147200 996147200 996147200 996147200' alert_rule_test: - alertname: MemoryUsage90Percent eval_time: 5m exp_alerts: - exp_labels: alertname: MemoryUsage90Percent severity: Critical alert_owner: alerts-infra container_label_io_rancher_stack_name: network-service instance: test-instance dns_name: test-dns exp_annotations: description: "Service memory has been at over 90% for 1 minute. Service: network-service, DnsName: test-dns."
内容的提问来源于stack exchange,提问作者Rad4
相关产品推荐
相关产品推荐

