You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取跨集群Informer Watch操作延迟的Prometheus指标?

监控Kubernetes Watch连接健康状态的方案

一、利用client-go内置Metrics

client-go本身提供了Watch相关的监控指标,无需从零实现:

  • 导入k8s.io/client-go/tools/metrics包,初始化时注册Watch指标:
    import "k8s.io/client-go/tools/metrics"
    
    func init() {
        // 注册默认Watch指标,包含请求延迟、断开次数等
        metrics.RegisterWatchMetrics(metrics.DefaultWatchMetrics())
    }
    
  • 内置核心指标:
    • watch_request_latency_seconds:Watch请求的建立延迟
    • watch_disconnects_total:Watch连接断开的总次数
    • watch_events_total:收到的Watch事件总数
  • 确保通过prometheus/client_golang暴露/metrics端点,让Prometheus完成指标采集

二、自定义Watch连接监控

如果内置指标无法满足需求,可通过包装client-go传输层或Watch接口实现细粒度监控:

1. 拦截Watch请求统计延迟

通过rest.Config的WrapTransport拦截请求,针对Watch请求单独统计:

import (
    "net/http"
    "time"

    "k8s.io/client-go/rest"
    "github.com/prometheus/client_golang/prometheus"
)

// 自定义Prometheus指标
var (
    watchLatency = prometheus.NewHistogramVec(
        prometheus.HistogramOpts{
            Name: "custom_watch_request_latency_seconds",
            Help: "Latency of watch requests to target cluster",
        },
        []string{"resource", "namespace"},
    )
    watchDisconnects = prometheus.NewCounterVec(
        prometheus.CounterOpts{
            Name: "custom_watch_disconnects_total",
            Help: "Total number of watch connections disconnected",
        },
        []string{"resource", "namespace"},
    )
)

func init() {
    prometheus.MustRegister(watchLatency, watchDisconnects)
}

func wrapWatchTransport(cfg *rest.Config) {
    original := cfg.WrapTransport
    cfg.WrapTransport = func(rt http.RoundTripper) http.RoundTripper {
        if original != nil {
            rt = original(rt)
        }
        return &watchMonitorRT{rt: rt}
    }
}

type watchMonitorRT struct {
    rt http.RoundTripper
}

func (w *watchMonitorRT) RoundTrip(req *http.Request) (*http.Response, error) {
    if req.URL.Query().Get("watch") != "true" {
        return w.rt.RoundTrip(req)
    }

    start := time.Now()
    resp, err := w.rt.RoundTrip(req)
    latency := time.Since(start).Seconds()

    // 解析请求的资源和命名空间
    resource := req.URL.Path
    ns := req.URL.Query().Get("namespace")
    watchLatency.WithLabelValues(resource, ns).Observe(latency)

    // 监听连接断开事件
    if resp != nil {
        if done, ok := resp.Body.(interface{ Done() <-chan struct{} }); ok {
            go func() {
                <-done.Done()
                watchDisconnects.WithLabelValues(resource, ns).Inc()
            }()
        }
    }

    return resp, err
}

2. 监听Watch接口的Stop信号

创建Watch实例后,直接监听其StopChan()捕获连接断开事件:

import (
    "context"
    "time"

    metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
    "k8s.io/client-go/kubernetes"
)

func monitorWatchConnection(clientset *kubernetes.Clientset, ns string) {
    watcher, err := clientset.CoreV1().Pods(ns).Watch(context.TODO(), metav1.ListOptions{})
    if err != nil {
        // 处理创建失败,记录错误指标
        return
    }
    defer watcher.Stop()

    connectTime := time.Now()
    <-watcher.StopChan()
    // 统计连接持续时长并暴露为指标
    duration := time.Since(connectTime).Seconds()
    // 示例:用Gauge记录当前连接持续时间,或Histogram统计断开时长
}

三、替代方案:监控目标集群连通性

若无需严格绑定Watch延迟,可定期发送轻量API请求作为连通性补充:

  • 比如定期请求/api/v1/namespaces/<target-ns>/pods?limit=1,统计请求延迟
  • 优点是实现简单,快速反映集群可达性;缺点是无法精确对应Watch连接状态

四、避坑提醒

  • 不要依赖Informer事件处理器监控Watch健康:事件触发依赖资源变更或重同步,缩短间隔会增加集群负载,不可取
  • 避免反射client-go内部结构体:依赖私有字段会导致代码在版本更新时崩溃,优先使用官方暴露的API或Metrics

内容的提问来源于stack exchange,提问作者helloSekai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 13:33:15