Prometheus指标重复采集错误排查与修复求助
问题
我用Go开发了一个提供RESTful API的服务,基于Fiber框架构建,通过中间件用Prometheus的CounterVec和Histogram类型指标(requests_total和request_duration)统计请求次数与耗时。服务部署在Kubernetes的Pod中,为了让每个Pod的指标唯一,给指标名称追加了GUID,实现参考Prometheus官方基础示例,使用的是client_golang v1.15.0库。
现在服务会偶发出现指标重复采集错误:启动运行一段时间后,抛出was collected before with the same name and label values错误,重试后最终崩溃,停止生成指标。我推测是竞态条件问题,但不知道怎么调试和用单元测试复现,而且给指标操作函数加了Mutex锁后问题依然存在,求排查、修复及复现方案。
错误信息如下:
4 error(s) occurred:
- collected metric "omittedPart_request_total_6687cdb68f_2wqht" { label:<name:"method" value:"POS" > label:<name:"path" value:"/api/someService/user/:user_id" > label:<name:"servicename" value:"some-service" > label:<name:"status_code" value:"204" > counter:<value:1 > } was collected before with the same name and label values
- collected metric "omittedPart_request_total_6687cdb68f_2wqht" { label:<name:"method" value:"POST" > label:<name:"path" value:"/api/someService" > label:<name:"servicename" value:"some-service" > label:<name:"status_code" value:"200" > counter:<value:39 > } was collected before with the same name and label values
- collected metric "omittedPart_request_total_6687cdb68f_2wqht" { label:<name:"method" value:"GETT" > label:<name:"path" value:"/api/someService" > label:<name:"servicename" value:"some-service" > label:<name:"status_code" value:"200" > counter:<value:1 > } was collected before with the same name and label values
- collected metric "omittedPart_request_duration_seconds_6687cdb68f_2wqht" { label:<name:"method" value:"POST" > label:<name:"path" value:"/api/ujt" > label:<name:"servicename" value:"some-service" > label:<name:"status_code" value:"200" > histogram:<sample_count:39 sample_sum:0.38108896699999995 bucket:<cumulative_count:0 upper_bound:1e-09 > bucket:<cumulative_count:0 upper_bound:2e-09 > bucket:<cumulative_count:0 upper_bound:5e-09 > bucket:<cumulative_count:0 upper_bound:1e-08 > bucket:<cumulative_count:0 upper_bound:2e-08 > bucket:<cumulative_count:0 upper_bound:5e-08 > bucket:<cumulative_count:0 upper_bound:1e-07 > bucket:<cumulative_count:0 upper_bound:2e-07 > bucket:<cumulative_count:0 upper_bound:5e-07 > bucket:<cumulative_count:0 upper_bound:1e-06 > bucket:<cumulative_count:0 upper_bound:2e-06 > bucket:<cumulative_count:0 upper_bound:5e-06 > bucket:<cumulative_count:0 upper_bound:1e-05 > bucket:<cumulative_count:0 upper_bound:2e-05 > bucket:<cumulative_count:0 upper_bound:5e-05 > bucket:<cumulative_count:0 upper_bound:0.0001 > bucket:<cumulative_count:0 upper_bound:0.0002 > bucket:<cumulative_count:0 upper_bound:0.0005 > bucket:<cumulative_count:0 upper_bound:0.001 > bucket:<cumulative_count:0 upper_bound:0.002 > bucket:<cumulative_count:0 upper_bound:0.005 > bucket:<cumulative_count:33 upper_bound:0.01 > bucket:<cumulative_count:38 upper_bound:0.02 > bucket:<cumulative_count:38 upper_bound:0.05 > bucket:<cumulative_count:39 upper_bound:0.1 > bucket:<cumulative_count:39 upper_bound:0.2 > bucket:<cumulative_count:39 upper_bound:0.5 > bucket:<cumulative_count:39 upper_bound:1 > bucket:<cumulative_count:39 upper_bound:2 > bucket:<cumulative_count:39 upper_bound:5 > bucket:<cumulative_count:39 upper_bound:10 > bucket:<cumulative_count:39 upper_bound:15 > bucket:<cumulative_count:39 upper_bound:20 > bucket:<cumulative_count:39 upper_bound:30 > > } was collected before with the same name and label values
排查方案
- 检查指标初始化逻辑:确认是否在请求链路(比如Fiber中间件)中重复创建
CounterVec/Histogram实例,或者GUID生成逻辑存在重复(概率低但需验证)。 - 验证标签值规范性:错误里出现
POS、GETT这类异常HTTP方法,说明标签值生成有bug——比如请求方法解析时被篡改、截断,导致同一实际请求生成异常标签,后续又重复注册相同标签组合的指标。 - 检查注册器使用:是否同时用了多个
prometheus.Registry实例,或者重复向同一个注册器注册同一指标(比如中间件初始化时多次执行注册逻辑)。 - 竞态检测:用
go build -race编译服务,开启竞态检测,同时模拟高并发请求,观察是否有数据竞争提示;在指标创建、注册、更新节点加日志,记录实例内存地址、标签值、注册时间,定位重复注册来源。
修复方案
- 确保指标实例全局唯一:把
CounterVec和Histogram定义为全局变量,仅在服务启动时初始化一次,避免在中间件或请求处理中重复创建:var ( requestsTotal *prometheus.CounterVec requestDuration *prometheus.HistogramVec ) func initMetrics(guid string) error { requestsTotal = prometheus.NewCounterVec( prometheus.CounterOpts{ Name: fmt.Sprintf("requests_total_%s", guid), Help: "Total number of requests", }, []string{"method", "path", "servicename", "status_code"}, ) requestDuration = prometheus.NewHistogramVec( prometheus.HistogramOpts{ Name: fmt.Sprintf("request_duration_seconds_%s", guid), Help: "Duration of requests in seconds", Buckets: prometheus.DefBuckets, }, []string{"method", "path", "servicename", "status_code"}, ) prometheus.MustRegister(requestsTotal) prometheus.MustRegister(requestDuration) return nil } - 修正标签值生成逻辑:检查Fiber中
c.Method()的使用是否被篡改,或者是否有大小写转换、字符串截断的bug,确保标签值为规范HTTP方法(如GET、POST)。 - 放弃自定义Mutex锁:
CounterVec和HistogramVec的WithLabelValues、Inc()、Observe()方法本身就是线程安全的,自定义锁反而可能导致逻辑混乱,直接使用官方方法即可。 - 避免重复注册:注册前判断指标是否已存在,或者复用已注册的实例:
if err := prometheus.Register(requestsTotal); err != nil { if are, ok := err.(prometheus.AlreadyRegisteredError); ok { requestsTotal = are.ExistingCollector.(*prometheus.CounterVec) } else { return err } }
复现方案
- 单元测试复现:模拟高并发下重复初始化指标,触发重复注册:
func TestMetricDuplicate(t *testing.T) { guid := "test-guid" // 多协程重复初始化指标 for i := 0; i < 10; i++ { go func() { err := initMetrics(guid) if err != nil && !strings.Contains(err.Error(), "already registered") { t.Fatal(err) } // 模拟指标更新 requestsTotal.WithLabelValues("POST", "/api/test", "test-service", "200").Inc() requestDuration.WithLabelValues("POST", "/api/test", "test-service", "200").Observe(0.1) }() } time.Sleep(1 * time.Second) // 采集指标并检查重复 registry := prometheus.DefaultRegisterer.(*prometheus.Registry) metrics, err := registry.Gather() if err != nil { t.Fatal(err) } metricCount := make(map[string]int) for _, m := range metrics { metricCount[*m.Name]++ if metricCount[*m.Name] > 1 { t.Fatalf("metric %s duplicated", *m.Name) } } } - 高并发压测复现:用
hey或wrk构造异常请求,同时压测服务:hey -n 10000 -c 100 -m POS http://localhost:3000/api/someService/user/123 - 手动模拟重复注册:在代码中创建两个相同名称、标签的
CounterVec实例,注册到同一注册器,验证是否触发相同错误。
内容的提问来源于stack exchange,提问作者rolldawg
相关产品推荐
相关产品推荐

