Kubernetes HPA自定义指标当前值异常:与Prometheus实际值不符
Hey there, let's break down why your HPA's metric values don't match what you're seeing in Prometheus. This is a common gotcha with custom metrics and HPA, so let's go through the most likely causes step by step:
1. Check Unit Conversion Discrepancies
First, let's decode that 132500m value in your HPA output: the m suffix stands for milli-units, so that translates to 132.5 (since 1 = 1000m). If your Prometheus query returns values way lower (like <100), this points to a unit mismatch in your metric pipeline:
- If you're using Prometheus Adapter to expose metrics to Kubernetes, double-check its config map. Did you accidentally apply a multiplier (e.g., multiplying raw values by 1000) when defining the custom metric?
- Verify the unit your exporter outputs. For example, some exporters emit metrics in "per second" units, while HPA might expect a different base unit.
2. Validate the HPA's Metric Query Logic
Your HPA pulls metrics from the Kubernetes custom metrics API (powered by Prometheus Adapter), so the query it uses might not match what you're running in Prometheus:
- First, get the exact metric definition used by your HPA:
Look for thekubectl describe hpa frontend -n streaming --context testingMetricssection—it will show the underlying PromQL (or metric name) the HPA is using. - Run that exact query in your Prometheus UI. If the result doesn't match what HPA shows, you've found the issue:
- Maybe the HPA is using
sum()over all pods, while you're checkingavg()in Prometheus? - Or the aggregation labels (like
pod,deployment) don't align between your manual query and the Adapter's config.
- Maybe the HPA is using
3. Account for Metric Caching & Latency
Kubernetes HPA relies on cached metrics from the custom metrics API, and Prometheus itself has a scrape interval:
- Prometheus Adapter typically caches metric results for ~30 seconds. Wait a minute or two, then recheck
kubectl get hpa—the value might sync up with Prometheus. - If the mismatch persists, check your Prometheus scrape interval. If it's set to 5 minutes, HPA will be working with stale data. Adjust the scrape interval for your target pods to something more frequent (e.g., 15 seconds) if needed.
4. Fix Namespace & Target Resource Misalignment
I noticed a potential red flag here: you ran kubectl get hpa -n frontend, but your HPA manifest specifies namespace: streaming. Make sure you're checking the right HPA:
- Run
kubectl get hpa -n streaming --context testingto confirm you're looking at the correct resource. - Also, verify the
scaleTargetRefin your HPA manifest points to the correct Deployment name. If it's targeting the wrong set of pods, the metrics will be unrelated to what you're checking in Prometheus.
5. Check Custom Metrics API Health
If the custom metrics API is broken, HPA might be returning stale or incorrect values:
- Verify the API service is available:
Thekubectl get apiservices | grep custom.metricsv1beta1.custom.metrics.k8s.ioentry should showSTATUS: Available. - If it's not, check the Prometheus Adapter logs for errors (replace
<adapter-pod>with your actual pod name):
Look for issues like failed PromQL queries, permission errors accessing Prometheus, or invalid metric configurations.kubectl logs -n monitoring <adapter-pod>
内容的提问来源于stack exchange,提问作者d.ansimov

