Telegraf DaemonSet采集K8s Pod指标遇Kafka连接异常求助
First, Let's Clarify the Metrics Question
When you run Telegraf as a DaemonSet in Kubernetes:
- By default (with basic inputs like
cpu,disk, etc.), it collects physical node metrics—this is because it’s mounted to the host’s/proc,/sys, and Docker socket, so it reads data directly from the node. - But with the
[[inputs.kubernetes]]plugin you’ve added in your config, it can absolutely collect Pod, container, and other Kubernetes-specific metrics—you just need to make sure the plugin has proper permissions to access the Kubernetes API.
Now let’s tackle your core issues: the Kafka connection failure and missing Pod metrics.
Step 1: Fix the Kafka Connection Error
The error kafka: client has run out of available brokers to talk to tells us the Telegraf pods can’t reliably reach your Kafka cluster, even though tcpdump shows some traffic. Here’s how to debug and fix this:
1.1 Verify Network Reachability from Telegraf Pods
First, confirm if the Telegraf pods can actually connect to your Kafka brokers:
- Exec into one of your Telegraf pods:
kubectl exec -n monitoring -it telegraf-dvtcl -- /bin/sh - Install
netcat(if missing) and test a broker’s port:
If this fails (but your node’s standalone Telegraf works), check:apt update && apt install -y netcat nc -zv 10.121.63.5 9092- Whether your Kubernetes CNI (Flannel, in your setup) has network policies restricting pod outbound traffic to Kafka’s 9092 port.
- If Kafka brokers have firewall rules blocking traffic from the pod’s network CIDR (not just the node’s IP).
1.2 Validate Kafka Output Configuration
Your Telegraf config has a few settings that might cause compatibility issues:
- Kafka Version: You set
version = "0.11.0.2", but Telegraf 1.9.2 uses Sarama 1.18.0. If your Kafka cluster runs a newer version (e.g., 2.x+), this mismatch can break connections. Update the version to match your cluster (e.g.,"2.0.0"for Kafka 2.x). - Compression Codec:
compression_codec = 2uses Snappy compression. Confirm your Kafka cluster has Snappy support enabled; if not, set this to0(no compression) to test. - Required Acks: Temporarily set
required_acks = 0(no acknowledgment needed) to rule out acknowledgment-related failures.
1.3 Match Your Working Standalone Telegraf Config
Since your node’s RPM Telegraf works perfectly, copy its Kafka output settings directly into your ConfigMap. This eliminates any configuration mismatches between the two deployments.
Step 2: Ensure Pod Metrics Are Collected
Even if Kafka is fixed, your inputs.kubernetes plugin might fail to fetch Pod metrics due to missing Kubernetes API permissions. Here’s how to fix that:
2.1 Create RBAC Permissions for Telegraf
The default ServiceAccount in the monitoring namespace doesn’t have permission to read Kubernetes resources like Pods. Create these RBAC resources:
apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole metadata: name: telegraf rules: - apiGroups: [""] resources: - pods - nodes - nodes/proxy - services - endpoints - namespaces verbs: ["get", "list", "watch"] --- apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRoleBinding metadata: name: telegraf subjects: - kind: ServiceAccount name: default namespace: monitoring roleRef: kind: ClusterRole name: telegraf apiGroup: rbac.authorization.k8s.io
Apply this with:
kubectl apply -f rbac-telegraf.yaml
2.2 Enable Debug Logging to Validate Metrics Collection
Update your ConfigMap to enable debug logs, which will show you if the kubernetes input is actually collecting metrics:
[agent] debug = true quiet = false
Restart the DaemonSet pods to apply changes:
kubectl rollout restart daemonset telegraf -n monitoring
Check the logs again with:
kubectl logs -n monitoring telegraf-dvtcl -f
Look for lines about collecting Kubernetes metrics—this will confirm if the plugin is working as expected.
Step 3: Verify End-to-End Data Flow
Once you’ve fixed the Kafka connection and RBAC permissions:
- Confirm Telegraf pods no longer throw Kafka connection errors.
- Run
telegraf --testinside a pod to see all collected metrics—this will let you verify Pod metrics are present. - Check your Kafka topic
telegraf-measurements-jsonto confirm data is being written (usekafka-console-consumer.shif you have access to the Kafka cluster).
内容的提问来源于stack exchange,提问作者Badb0y

