使用Prometheus Go SDK向Pushgateway推送Histogram指标后仅显示+Inf分桶的问题求助
各位大佬,我遇到一个用Prometheus Go SDK推送Histogram指标到Pushgateway的棘手问题,折腾好久没搞定,来求助大家!
问题背景
我用Go的Prometheus SDK生成Histogram类型的指标,特意设置了多个指数分桶,然后推送到Pushgateway。调试时已经确认请求体里包含所有分桶的数据,但在Grafana里却只能看到+Inf的那个分桶,其他分桶完全看不到。更奇怪的是,用Python的prometheus_client写了个类似的推送脚本,就能正常显示所有分桶,所以我怀疑是不是Go SDK的用法哪里出问题了?
预期结果
我期望在Grafana里能看到所有定义的分桶数据,比如这样:
runoncejob_nolive_execution_launch_duration_seconds_bucket{cid="sg",env="live",execution_type="adhoc",group_name="runonce",group_namespace="runonce",job_id="job_test_0",le="0.5"} 0
runoncejob_nolive_execution_launch_duration_seconds_bucket{cid="sg",env="live",execution_type="adhoc",group_name="runonce",group_namespace="runonce",job_id="job_test_0",le="1"} 0
runoncejob_nolive_execution_launch_duration_seconds_bucket{cid="sg",env="live",execution_type="adhoc",group_name="runonce",group_namespace="runonce",job_id="job_test_0",le="2"} 0
runoncejob_nolive_execution_launch_duration_seconds_bucket{cid="sg",env="live",execution_type="adhoc",group_name="runonce",group_namespace="runonce",job_id="job_test_0",le="+Inf"} 1
实际结果
在Grafana里只能看到标记为le="+Inf"的分桶,其他预先定义的分桶完全找不到踪影。
我使用的Go代码
func TestPush(t *testing.T) { jobLevelExecutionLaunchDuration := prometheus.NewHistogramVec( prometheus.HistogramOpts{ Namespace: "runoncejob", Name: "nolive_execution_launch_duration_seconds", Help: "Duration that executions spent waiting for the recent task to be launched, " + "observed once task is launched.", Buckets: prometheus.ExponentialBuckets(0.5, 2, 3), }, getJobLevelExecutionLabels("execution_type"), ) for i := 0; i < 1; i++ { sleep := time.Duration(rand.Intn(4000)) * time.Millisecond time.Sleep(sleep) value := float64(rand.Intn(400)) jobId := fmt.Sprintf("job_test_%d", i%5) jobLevelExecutionLaunchDuration.WithLabelValues("runonce", "runonce", jobId, "live", "sg", "adhoc").Observe(value) } reg := prometheus.NewRegistry() reg.MustRegister(jobLevelExecutionLaunchDuration) pusher := push.New("xxxxxxxxxxxx", "test_job") if err := pusher.Gatherer(reg).Push(); err != nil { fmt.Println("Could not push completion time to Pushgateway:", err) } fmt.Println("Pushed completion time to Pushgateway") } func getJobLevelExecutionLabels(labels ...string) []string { return metrics.AppendStrings([]string{ "group_namespace", "group_name", "job_id", "env", "cid", }, labels...) }
已做的排查动作
- 抓包+调试确认:推送给Pushgateway的请求体里确实包含了所有分桶的数据;
- 用Python的prometheus_client写了对比脚本,推送后能在Grafana正常看到所有分桶,Python代码如下:
def run_0630(): push_url = "xxxxxxxx" push_job = "test_job" push_interval = 15 registry = CollectorRegistry() test_metric = Histogram( "runoncejob_nolive_execution_launch_duration_seconds", "testing metric", ["cluster"], buckets=[0.1, 0.2, 0.5, 1.0, 5.0, 10.0], registry=registry, ) while True: test_metric.labels( cluster = "test", ).observe(10) print("pushing metrics") push_to_gateway(push_url, job=push_job, registry=registry) print("pushed metrics") time.sleep(push_interval)
有没有大佬能帮我看看Go代码里哪里写错了?是不是HistogramVec的注册或者推送方式有什么细节我没注意到?
内容来源于stack exchange

