Prometheus Targets Errors排查:K8s集群节点Prometheus信息获取问题
Hey there! Let's break down how to troubleshoot and fix those Prometheus Targets errors you're hitting when trying to sync metrics to the rest of your Kubernetes nodes.
Before jumping into fixes, you need to know what exactly is failing. Here's how to get the specifics:
- Fire up the Prometheus UI: Use
kubectl port-forward <prometheus-pod-name> 9090:9090to access it locally, then navigate to Status > Targets. Look at theLast Errorcolumn for each target—this will tell you if it's a connection timeout, permission issue, or config problem. - Check Prometheus logs: Run
kubectl logs <prometheus-pod-name> -n <your-monitoring-namespace>and search for phrases liketarget scrape failedor specific error messages to get context.
Let's go through the most frequent issues and how to resolve them:
1. Connection Timeouts / Target Unreachable
This is one of the most common culprits—Prometheus can't reach the metrics endpoint on your nodes (usually node-exporter on port 9100).
- Check node-exporter deployment: Run
kubectl get daemonsets -n <monitoring-namespace>to make sure the node-exporter DaemonSet has all pods running (DESIRED should match READY). - Verify network access: If your cluster uses network policies, ensure the namespace where Prometheus runs has permission to access port 9100 on all nodes. You can test connectivity directly from the Prometheus pod:
If this fails, you know it's a network-level issue.kubectl exec -it <prometheus-pod-name> -n <namespace> -- curl <target-node-ip>:9100/metrics
2. Permission Denied / Authentication Errors
If Prometheus gets a "permission denied" or 401 error, it's missing credentials to access the metrics endpoint.
- Basic auth setup: If your node-exporter uses basic auth, add the credentials to your Prometheus scrape config:
scrape_configs: - job_name: 'node-exporter' basic_auth: username: 'your-metrics-username' password: 'your-metrics-password' static_configs: - targets: ['node-ip-1:9100', 'node-ip-2:9100'] - RBAC permissions: Ensure Prometheus's service account has the right roles to discover and access node resources. You can bind a
viewcluster role to it, or create a custom role with permissions for node metrics.
3. Invalid Scrape Configuration
Sometimes the issue is a typo or misconfiguration in how Prometheus discovers targets.
- Check K8s service discovery config: If you're using Kubernetes service discovery for nodes, make sure your scrape job looks something like this (note the
role: nodeand relabeling):scrape_configs: - job_name: 'kubernetes-nodes' kubernetes_sd_configs: - role: node relabel_configs: - action: labelmap regex: __meta_kubernetes_node_label_(.+) - target_label: __address__ replacement: '${1}:9100' - Reload Prometheus config: After updating your ConfigMap, restart the Prometheus deployment with
kubectl rollout restart deployment <prometheus-deployment> -n <namespace>, or use the UI's Configuration > Reload button.
4. Node-Exporter Isn't Exposing Metrics Correctly
If node-exporter isn't set up right, Prometheus can't scrape anything.
- Check listen address: Make sure node-exporter is listening on all interfaces (not just localhost) with the flag
--web.listen-address=0.0.0.0:9100. You can verify this in the pod's startup command. - Check node-exporter logs: Run
kubectl logs <node-exporter-pod-name> -n <namespace>to look for startup errors or issues binding to the port.
Once you've applied a fix, head back to the Prometheus Targets page—you should see the target status switch to UP within a minute or two. You can also confirm with a query in the Graph tab: run up{job="node-exporter"} to check if all nodes are reporting as up.
内容的提问来源于stack exchange,提问作者Alejandro López

