基于请求TPS的Kubernetes HPA配置故障:Prometheus自定义指标无法被识别
Let's break down why you're hitting the NotFound error and walk through step-by-step fixes:
1. Mismatched seriesQuery Configuration
Your seriesQuery is set to look for metrics matching dapr_http_server_request_avg_.*, but the actual metric you want to use is dapr_http_server_request_count. Prometheus Adapter relies on seriesQuery to discover existing time series in Prometheus—if this query returns no results, the adapter can't create the corresponding custom Kubernetes metric at all.
Fix:
Update the seriesQuery to target your actual metric, and include labels that link to Kubernetes pod/namespace resources:
seriesQuery: '{__name__="dapr_http_server_request_count",app_id="governor",path=~"/v1.0/invoke/app/method/interceptor/.*",namespace!="",pod!=""}'
2. Incorrect name Matching Rule
Your name.matches pattern targets metrics ending in _total, but your metric ends with _count. This mismatch means the adapter can't properly rename the metric to your desired _per_second format, so it won't register the custom metric.
Fix:
Adjust the matches pattern to fit your metric's suffix, while keeping the as rule for a TPS-friendly name:
name: matches: "^(.*)_count" as: "${1}_per_second"
3. metricsQuery Loses Pod/Namespace Context
Your current metricsQuery uses a double sum that strips away all pod and namespace labels:
sum(sum by (path) (rate(dapr_http_server_request_count{app_id="governor",path=~"/v1.0/invoke/app/method/interceptor/.*"}[10s])))
HPA needs metrics tied to individual pods (or linked to pod/namespace resources) to scale correctly. If you aggregate away these labels, the adapter can't map the metric to your Kubernetes pods.
Fix:
Use the <<.Series>> placeholder to reference the series found by seriesQuery, and preserve the critical pod and namespace labels in your aggregation:
metricsQuery: 'sum by (pod, namespace) (sum by (pod, namespace, path) (rate(<<.Series>>[10s])))'
This first aggregates by path (to sum all interceptor paths per pod) then sums those values per pod/namespace—keeping the labels the adapter needs to link the metric to your pods.
4. Wrong Metric Name in kubectl Query
You're querying for dapr_http_server_request_avg_total, but based on the corrected name rules, your custom metric should be dapr_http_server_request_per_second. Using the wrong name will always result in a NotFound error.
Fix:
Use the correct metric name in your kubectl command:
kubectl get --raw "/apis/custom.metrics.k8s.io/v1beta1/namespaces/app/pods/*/dapr_http_server_request_per_second" | jq .
Additional Validation Steps
- Test the PromQL in Prometheus UI: First run your corrected metric query directly in Prometheus to confirm it returns results with
podandnamespacelabels. - Check Prometheus Adapter Logs: Look for errors in the adapter's logs to confirm it's discovering metrics correctly:
kubectl logs -n <your-adapter-namespace> <prometheus-adapter-pod-name> - Restart the Adapter: After updating the configmap, restart the Prometheus Adapter pods to load the new configuration.
Sample Corrected Prometheus Adapter Config
Putting it all together, your custom config should look like this:
default: true custom: - seriesQuery: '{__name__="dapr_http_server_request_count",app_id="governor",path=~"/v1.0/invoke/app/method/interceptor/.*",namespace!="",pod!=""}' resources: overrides: namespace: resource: namespace pod: resource: pod name: matches: "^(.*)_count" as: "${1}_per_second" metricsQuery: 'sum by (pod, namespace) (sum by (pod, namespace, path) (rate(<<.Series>>[10s])))'
内容的提问来源于stack exchange,提问作者Uday Chauhan

