You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

自定义K8s指标采集器:如何暴露/metrics端点并实现循环采集

自定义K8s指标采集器:暴露/metrics路径与跨集群版本监控方案

一、暴露/metrics路径的实现

使用prometheus_client库的start_http_server方法启动HTTP服务器后,默认会在指定端口暴露/metrics路径,你当前代码中的start_http_server(8000)已经完成了这个工作。启动后,访问http://<你的采集器地址>:8000/metrics即可获取指标数据,Prometheus可直接配置该地址作为scrape目标。

二、无限循环采集的逻辑说明

不需要在代码中主动编写采集循环,因为Prometheus是定时拉取/metrics路径的。每次拉取请求触发时,CustomCollector的collect方法会自动执行,重新采集最新的K8s数据并生成指标。代码末尾的while True循环仅用于保持采集器进程持续运行,避免退出。

三、跨集群监控与代码优化

针对你监控两个集群应用版本的需求,以及原代码中的问题(如重复输出指标、未支持多集群),以下是修正后的完整代码:

import time
from prometheus_client import start_http_server
from prometheus_client.core import REGISTRY, CounterMetricFamily
from kubernetes import client, config

class CustomCollector(object):
    def __init__(self):
        # 初始化两个集群的API客户端,替换为你的kubeconfig路径
        self.cluster_clients = self._setup_multicluster_clients()

    def _setup_multicluster_clients(self):
        """加载多个集群的kubeconfig,返回集群API客户端列表"""
        clients = []
        # 集群1配置
        cluster1_config = config.load_kube_config('config-cluster1')
        clients.append({
            'cluster_name': 'cluster-1',
            'core_api': client.CoreV1Api(cluster1_config),
            'custom_api': client.CustomObjectsApi(client.ApiClient(cluster1_config))
        })
        # 集群2配置
        cluster2_config = config.load_kube_config('config-cluster2')
        clients.append({
            'cluster_name': 'cluster-2',
            'core_api': client.CoreV1Api(cluster2_config),
            'custom_api': client.CustomObjectsApi(client.ApiClient(cluster2_config))
        })
        return clients

    def collect(self):
        # 定义指标,新增cluster标签区分不同集群
        pod_info_metric = CounterMetricFamily(
            "retail_pods_info",
            "Pod information linked with Argo CD application version",
            labels=['cluster', 'secret', 'namespace', 'deployment_name', 'image', 'helm_version']
        )
        # Argo CD自定义资源配置
        argocd_group = "argoproj.io"
        argocd_version = "v1alpha1"
        argocd_plural = "applications"
        argocd_namespace = "argo-cd"

        # 遍历每个集群采集数据
        for cluster in self.cluster_clients:
            cluster_name = cluster['cluster_name']
            core_api = cluster['core_api']
            custom_api = cluster['custom_api']

            # 采集所有Pod的相关信息
            pod_list = core_api.list_pod_for_all_namespaces(watch=False)
            pod_metrics = []
            for pod in pod_list.items:
                metadata = pod.metadata
                spec = pod.spec
                if not spec.volumes:
                    continue
                # 筛选包含Secret的Projected Volume
                for volume in spec.volumes:
                    if volume.projected:
                        for source in volume.projected.sources:
                            if source.secret:
                                secret_name = source.secret.name
                                namespace = metadata.namespace.lower()
                                # 从Pod名称提取Deployment名称(格式:deployment-name-xxx-xxx)
                                deployment_name = metadata.name.lower().rsplit('-', 2)[0]
                                container_image = pod.spec.containers[0].image
                                pod_metrics.append([secret_name, namespace, deployment_name, container_image])

            # 获取Argo CD应用列表
            argocd_apps = custom_api.list_namespaced_custom_object(
                argocd_group,
                argocd_version,
                argocd_namespace,
                argocd_plural,
                watch=False
            )

            # 关联Pod与Argo CD应用版本
            for metric in pod_metrics:
                secret, ns, deploy_name, image = metric
                for app in argocd_apps["items"]:
                    if deploy_name == app["metadata"]["name"]:
                        helm_version = f"{app['spec']['source']['repoURL']}-{app['spec']['source']['targetRevision']}"
                        pod_info_metric.add_metric([cluster_name, secret, ns, deploy_name, image, helm_version], 1)
        # 统一输出所有指标,避免重复生成
        yield pod_info_metric

if __name__ == '__main__':
    # 启动HTTP服务器,监听8000端口,默认暴露/metrics
    start_http_server(8000)
    # 注册自定义采集器
    REGISTRY.register(CustomCollector())
    # 保持进程运行,等待Prometheus拉取
    while True:
        time.sleep(3600)  # 长时间休眠,减少资源消耗

关键优化点

  • 多集群支持:通过加载不同集群的kubeconfig,初始化多个API客户端,实现跨集群指标采集
  • 指标正确性:统一在所有数据处理完成后yield指标,避免原代码中重复输出同一指标的问题
  • 逻辑简化:移除不必要的变量,优化循环嵌套逻辑,提升代码可读性
  • 标签增强:新增cluster标签,清晰区分不同集群的指标数据

内容的提问来源于stack exchange,提问作者Garamoff

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 08:13:12