You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GKE集群HPA无法获取StackDriver自定义指标问题排查与解决

解决GKE集群中基于StackDriver自定义指标的HPA无法获取指标问题

我最近在GKE集群里碰到了个棘手的问题:想通过上报到StackDriver的自定义指标num_drivers_per_pod实现Pod自动扩缩容,前期准备都到位了——指标能在StackDriver的Metric Explorer里查到,用kubectl get --raw "/apis/custom.metrics.k8s.io/v1beta1"也能看到这个自定义指标,StackDriver Adapter和Heapster也部署好了,但一部署HPA清单就报错:

unable to get metric num_drivers_per_pod: no metrics returned from custom metrics API

先给大家看看我最初用来上报指标的Python代码:

import os
import time
import traceback
import logging
from google.cloud import monitoring_v3

logger = logging.getLogger(__name__)

def put_k8_pod_metric(metric_name,value,metric_type="k8s_pod"):
    try:
        client = monitoring_v3.MetricServiceClient()
        series = monitoring_v3.types.TimeSeries()
        series.metric.type = f'custom.googleapis.com/{metric_name}'
        series.resource.type = metric_type
        series.resource.labels['project_id'] = os.getenv("PROJECT_NAME")
        series.resource.labels['location'] = os.getenv("POD_LOCATION","asia-south1")
        series.resource.labels['cluster_name'] = os.getenv("CLUSTER_NAME","data-k8cluster")
        series.resource.labels['namespace_name'] = "default"
        series.resource.labels['pod_name'] = os.getenv("MY_POD_NAME","wrong_pod")
        point = series.points.add()
        point.value.double_value = value
        now = time.time()
        point.interval.end_time.seconds = int(now)
        point.interval.end_time.nanos = int( (now - point.interval.end_time.seconds) * 10**9)
        project_name = client.project_path(os.getenv('PROJECT_NAME'))
        client.create_time_series(project_name, [series],timeout=2)
        logger.info(f"successfully send the metric {metric_name} with value {value}")
    except Exception as e:
        traceback.print_exc()
        logger.info(f"failed to send the metric {metric_name} with value {value}")

经过一番排查,终于找到问题根源,做了两个关键修改后问题就解决了:

  • 升级HPA的API版本:把HPA配置文件里的API版本从autoscaling/v1升级到autoscaling/v2beta2或者autoscaling/v2,旧版本API对自定义指标的支持不完善,无法和StackDriver Adapter正确交互。
  • 调整指标上报的资源类型:把代码里的metric_type默认值从k8s_pod改回gke_container。GKE环境下,StackDriver Adapter需要识别gke_container类型的资源才能正确关联到对应Pod,之前用k8s_pod导致指标和Pod的关联关系没被正确识别,HPA自然拿不到数据。

修改完这两处后,HPA就能正常获取num_drivers_per_pod指标,Pod的自动扩缩容也正常工作了。我已经把调整后的完整示例代码整理成了仓库,包含正确的指标上报逻辑和HPA配置。

内容的提问来源于stack exchange,提问作者navdeep

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 08:39:10