You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Prometheus Go SDK向Pushgateway推送Histogram指标后仅显示+Inf分桶的问题求助

Prometheus Go SDK向Pushgateway推送Histogram指标后仅显示+Inf分桶的问题求助

各位大佬,我遇到一个用Prometheus Go SDK推送Histogram指标到Pushgateway的棘手问题,折腾好久没搞定,来求助大家!

问题背景

我用Go的Prometheus SDK生成Histogram类型的指标,特意设置了多个指数分桶,然后推送到Pushgateway。调试时已经确认请求体里包含所有分桶的数据,但在Grafana里却只能看到+Inf的那个分桶,其他分桶完全看不到。更奇怪的是,用Python的prometheus_client写了个类似的推送脚本,就能正常显示所有分桶,所以我怀疑是不是Go SDK的用法哪里出问题了?

预期结果

我期望在Grafana里能看到所有定义的分桶数据,比如这样:

runoncejob_nolive_execution_launch_duration_seconds_bucket{cid="sg",env="live",execution_type="adhoc",group_name="runonce",group_namespace="runonce",job_id="job_test_0",le="0.5"} 0
runoncejob_nolive_execution_launch_duration_seconds_bucket{cid="sg",env="live",execution_type="adhoc",group_name="runonce",group_namespace="runonce",job_id="job_test_0",le="1"} 0
runoncejob_nolive_execution_launch_duration_seconds_bucket{cid="sg",env="live",execution_type="adhoc",group_name="runonce",group_namespace="runonce",job_id="job_test_0",le="2"} 0
runoncejob_nolive_execution_launch_duration_seconds_bucket{cid="sg",env="live",execution_type="adhoc",group_name="runonce",group_namespace="runonce",job_id="job_test_0",le="+Inf"} 1

实际结果

在Grafana里只能看到标记为le="+Inf"的分桶,其他预先定义的分桶完全找不到踪影。

我使用的Go代码

func TestPush(t *testing.T) {
    jobLevelExecutionLaunchDuration := prometheus.NewHistogramVec(
        prometheus.HistogramOpts{
            Namespace: "runoncejob",
            Name:      "nolive_execution_launch_duration_seconds",
            Help: "Duration that executions spent waiting for the recent task to be launched, " +
                "observed once task is launched.",
            Buckets: prometheus.ExponentialBuckets(0.5, 2, 3),
        },
        getJobLevelExecutionLabels("execution_type"),
    )

    for i := 0; i < 1; i++ {
        sleep := time.Duration(rand.Intn(4000)) * time.Millisecond
        time.Sleep(sleep)
        value := float64(rand.Intn(400))
        jobId := fmt.Sprintf("job_test_%d", i%5)
        jobLevelExecutionLaunchDuration.WithLabelValues("runonce", "runonce", jobId, "live", "sg", "adhoc").Observe(value)
    }

    reg := prometheus.NewRegistry()
    reg.MustRegister(jobLevelExecutionLaunchDuration)

    pusher := push.New("xxxxxxxxxxxx", "test_job")
    if err := pusher.Gatherer(reg).Push(); err != nil {
        fmt.Println("Could not push completion time to Pushgateway:", err)
    }
    fmt.Println("Pushed completion time to Pushgateway")
}

func getJobLevelExecutionLabels(labels ...string) []string {
    return metrics.AppendStrings([]string{
        "group_namespace",
        "group_name",
        "job_id",
        "env",
        "cid",
    }, labels...)
}

已做的排查动作

  • 抓包+调试确认:推送给Pushgateway的请求体里确实包含了所有分桶的数据;
  • 用Python的prometheus_client写了对比脚本,推送后能在Grafana正常看到所有分桶,Python代码如下:
def run_0630():
    push_url = "xxxxxxxx"
    push_job = "test_job"
    push_interval = 15
    registry = CollectorRegistry()
    test_metric = Histogram(
        "runoncejob_nolive_execution_launch_duration_seconds",
        "testing metric",
        ["cluster"],
        buckets=[0.1, 0.2, 0.5, 1.0, 5.0, 10.0],
        registry=registry,
    )
    while True:
        test_metric.labels(
            cluster = "test",
        ).observe(10)
        print("pushing metrics")
        push_to_gateway(push_url, job=push_job, registry=registry)
        print("pushed metrics")
        time.sleep(push_interval)

有没有大佬能帮我看看Go代码里哪里写错了?是不是HistogramVec的注册或者推送方式有什么细节我没注意到?

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.08 10:58:01