HPA无法获取MongoDB Exporter自定义指标问题排查求助
First off, let’s confirm the basics—you can query the metric directly via the custom metrics API, so the problem isn’t with the mongodb-exporter or custom metrics server itself. It’s almost certainly a mismatch between how your HPA is configured and how the metric is exposed. Here’s what to check step by step:
Verify exact metric name match
Custom metrics are case-sensitive and require perfect name matching. Copy themetricNamevalue directly from your API query response (runkubectl get --raw "/apis/custom.metrics.k8s.io/v1beta1/namespaces/monitoring/pods/*/mongodb_mongod_wiredtiger_cache_bytes" | jq '.items[].metricName'to extract it) and paste it into your HPA’sspec.metrics[0].pods.metric.namefield. Even a single missing underscore or capitalization error will break the HPA’s ability to find the metric.Ensure HPA targets the right pods
Your HPA’sscaleTargetRefmust point to the correct Deployment/StatefulSet in themonitoringnamespace. Additionally, check that the pods returned by your custom metrics API have matching labels to the pods managed by your target resource. Runkubectl get pods -n monitoring --show-labelsto list pod labels, then compare them to thelabelsfield in your API response. If your HPA includes aselectorfield, it must exactly match labels on the metric-bearing pods.Check kube-controller-manager permissions
The HPA controller (part of kube-controller-manager) needs permission to query the custom metrics API. Verify that the appropriate ClusterRole and ClusterRoleBinding exist to allow thesystem:kube-controller-managerservice account to accesscustom.metrics.k8s.ioresources. You can check the controller’s logs for permission errors with:kubectl logs -n kube-system <your-kube-controller-manager-pod> | grep -i "custom.metrics\|mongodb_mongod_wiredtiger_cache_bytes"Look for entries like "permission denied" or "unable to list metrics".
Confirm metric type compatibility
HPA’sutilizationtarget expects a gauge metric (a value that represents current state, like cache usage). If your metric is a counter (an ever-increasing value), usingutilizationwill cause issues. Check thetypefield in your API response—it should beGauge. If it’s a Counter, switch your HPA target to useAverageValueinstead ofUtilization, or ensure the mongodb-exporter is exposing the metric as a Gauge.Validate HPA API version
Older HPA API versions (likeautoscaling/v1) don’t support custom metrics. Make sure your HPA usesautoscaling/v2orautoscaling/v2beta2in theapiVersionfield. Here’s a working example template for your use case:apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: mongodb-cache-hpa namespace: monitoring spec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment # Or StatefulSet if you're using one name: mongodb-deployment # Replace with your target resource name minReplicas: 1 maxReplicas: 5 metrics: - type: Pods pods: metric: name: mongodb_mongod_wiredtiger_cache_bytes target: type: AverageValue # Use this if metric is a Gauge, or adjust as needed averageValue: 1Gi # Replace with your desired thresholdCheck metric server cache and refresh rates
Custom metrics servers often cache data. If your HPA queries before the metric server has synced the latest data from Prometheus, it might get an empty response. Check your metric server’s logs for sync errors, and ensure its refresh interval aligns with Prometheus’s scrape interval (default is 15s for Prometheus—set your metric server to sync at least that often).
If none of these fix the issue, dig deeper into the kube-controller-manager logs—they’ll usually include a more specific error message about why the metric isn’t being returned (e.g., no pods matching the selector, invalid metric type).
内容的提问来源于stack exchange,提问作者Ayhan Balik

