You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Telegraf DaemonSet采集K8s Pod指标遇Kafka连接异常求助

Telegraf DaemonSet in Kubernetes: Node vs Pod Metrics + Kafka Connection Fix

First, Let's Clarify the Metrics Question

When you run Telegraf as a DaemonSet in Kubernetes:

  • By default (with basic inputs like cpu, disk, etc.), it collects physical node metrics—this is because it’s mounted to the host’s /proc, /sys, and Docker socket, so it reads data directly from the node.
  • But with the [[inputs.kubernetes]] plugin you’ve added in your config, it can absolutely collect Pod, container, and other Kubernetes-specific metrics—you just need to make sure the plugin has proper permissions to access the Kubernetes API.

Now let’s tackle your core issues: the Kafka connection failure and missing Pod metrics.


Step 1: Fix the Kafka Connection Error

The error kafka: client has run out of available brokers to talk to tells us the Telegraf pods can’t reliably reach your Kafka cluster, even though tcpdump shows some traffic. Here’s how to debug and fix this:

1.1 Verify Network Reachability from Telegraf Pods

First, confirm if the Telegraf pods can actually connect to your Kafka brokers:

  • Exec into one of your Telegraf pods:
    kubectl exec -n monitoring -it telegraf-dvtcl -- /bin/sh
    
  • Install netcat (if missing) and test a broker’s port:
    apt update && apt install -y netcat
    nc -zv 10.121.63.5 9092
    
    If this fails (but your node’s standalone Telegraf works), check:
    • Whether your Kubernetes CNI (Flannel, in your setup) has network policies restricting pod outbound traffic to Kafka’s 9092 port.
    • If Kafka brokers have firewall rules blocking traffic from the pod’s network CIDR (not just the node’s IP).

1.2 Validate Kafka Output Configuration

Your Telegraf config has a few settings that might cause compatibility issues:

  • Kafka Version: You set version = "0.11.0.2", but Telegraf 1.9.2 uses Sarama 1.18.0. If your Kafka cluster runs a newer version (e.g., 2.x+), this mismatch can break connections. Update the version to match your cluster (e.g., "2.0.0" for Kafka 2.x).
  • Compression Codec: compression_codec = 2 uses Snappy compression. Confirm your Kafka cluster has Snappy support enabled; if not, set this to 0 (no compression) to test.
  • Required Acks: Temporarily set required_acks = 0 (no acknowledgment needed) to rule out acknowledgment-related failures.

1.3 Match Your Working Standalone Telegraf Config

Since your node’s RPM Telegraf works perfectly, copy its Kafka output settings directly into your ConfigMap. This eliminates any configuration mismatches between the two deployments.


Step 2: Ensure Pod Metrics Are Collected

Even if Kafka is fixed, your inputs.kubernetes plugin might fail to fetch Pod metrics due to missing Kubernetes API permissions. Here’s how to fix that:

2.1 Create RBAC Permissions for Telegraf

The default ServiceAccount in the monitoring namespace doesn’t have permission to read Kubernetes resources like Pods. Create these RBAC resources:

apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: telegraf
rules:
- apiGroups: [""]
  resources:
  - pods
  - nodes
  - nodes/proxy
  - services
  - endpoints
  - namespaces
  verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: telegraf
subjects:
- kind: ServiceAccount
  name: default
  namespace: monitoring
roleRef:
  kind: ClusterRole
  name: telegraf
  apiGroup: rbac.authorization.k8s.io

Apply this with:

kubectl apply -f rbac-telegraf.yaml

2.2 Enable Debug Logging to Validate Metrics Collection

Update your ConfigMap to enable debug logs, which will show you if the kubernetes input is actually collecting metrics:

[agent]
  debug = true
  quiet = false

Restart the DaemonSet pods to apply changes:

kubectl rollout restart daemonset telegraf -n monitoring

Check the logs again with:

kubectl logs -n monitoring telegraf-dvtcl -f

Look for lines about collecting Kubernetes metrics—this will confirm if the plugin is working as expected.


Step 3: Verify End-to-End Data Flow

Once you’ve fixed the Kafka connection and RBAC permissions:

  1. Confirm Telegraf pods no longer throw Kafka connection errors.
  2. Run telegraf --test inside a pod to see all collected metrics—this will let you verify Pod metrics are present.
  3. Check your Kafka topic telegraf-measurements-json to confirm data is being written (use kafka-console-consumer.sh if you have access to the Kafka cluster).

内容的提问来源于stack exchange,提问作者Badb0y

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:58:35