如何将GCP Kubernetes集群中的Pod添加至已部署的Zabbix监控系统?
Hey there! Since you've already got Zabbix Server and the Agent DaemonSet up on your GCP Kubernetes cluster, let's walk through how to add your running Pods to Zabbix monitoring. Here's a practical, step-by-step approach:
Before diving into Pod monitoring, make sure the Agent is active on every cluster node. Run this command to check:
kubectl get daemonsets -n <your-zabbix-namespace>
Look at the READY column—this number should match the total number of nodes in your cluster. If it doesn't, troubleshoot the DaemonSet first (check logs, node taints, etc.)
There are two reliable ways to monitor Kubernetes Pods with Zabbix; pick the one that fits your setup best:
Option 1: Kubernetes API Discovery (Recommended)
This method lets Zabbix automatically discover and monitor Pods using the Kubernetes API, which is scalable and low-maintenance:
- Create a Host Group: In the Zabbix Web UI, go to
Configuration > Host Groupsand make a new group (e.g.,Kubernetes Pods) to organize all your Pod monitoring entries. - Set Up a Pod Monitoring Template: Create a new template (or adapt an existing Kubernetes template) and add key monitoring items like:
- Pod status (tracks if a Pod is
Running,Failed, etc.) - CPU and memory usage per Pod
- Restart count for Pod containers
- Pod status (tracks if a Pod is
- Configure a Discovery Rule:
- Head to
Configuration > Discoveryin Zabbix and create a new rule. - Set the discovery type to
Kubernetes, then enter your cluster's internal API endpoint:https://kubernetes.default.svc. - For authentication, use a Kubernetes ServiceAccount with
viewpermissions (create one if you don't have it). Grab the ServiceAccount's token and cluster CA certificate, then paste them into the discovery rule's auth fields. - Add filters to target specific Pods (e.g., by namespace, label) or leave it open to monitor all Pods.
- Create an Action to auto-add discovered Pods to your
Kubernetes Podshost group and link them to your Pod monitoring template.
- Head to
Option 2: Custom Checks via Zabbix Agent
If you prefer using the node-level Agents to pull Pod data, you can add custom checks:
- Update the Agent ConfigMap: Edit the ConfigMap used by your Zabbix Agent DaemonSet to add these
UserParameterentries:# Get status of all Pods in a namespace UserParameter=kubernetes.pods.status[*],kubectl get pods -n $1 -o jsonpath='{range .items[*]}{.metadata.name}:{.status.phase}{"\n"}{end}' # Get CPU usage for a specific Pod UserParameter=kubernetes.pod.cpu[*],kubectl top pod $1 -n $2 --no-headers | awk '{print $2}' # Get memory usage for a specific Pod UserParameter=kubernetes.pod.memory[*],kubectl top pod $1 -n $2 --no-headers | awk '{print $3}' - Ensure Agent Has Kubectl & Permissions: Make sure your Agent Pod includes the
kubectltool (update the Agent image if needed) and mount the ServiceAccount token/CA cert to let it access the Kubernetes API. - Add Monitoring Items in Zabbix: In the Zabbix UI, create new monitoring items linked to your node Agents, using the custom parameters you defined. You can use macros or regex to target specific Pods across nodes.
- Wait 5-10 minutes for Zabbix to pull data, then go to
Monitoring > Hoststo check if your Pods appear in the host group. - Click into a Pod host to verify that monitoring items are collecting data (look for green "OK" statuses).
- If something's broken, check logs:
- Zabbix Server logs:
kubectl logs <zabbix-server-pod-name> -n <your-zabbix-namespace> - Zabbix Agent logs:
kubectl logs -l app=zabbix-agent -n <your-zabbix-namespace>
- Zabbix Server logs:
- Add Triggers: Set up alerts for critical events, like a Pod entering
CrashLoopBackOffor exceeding CPU/memory thresholds. - Build Graphs: Create visual graphs for Pod resource usage to spot trends quickly.
- Customize for Apps: For application-specific Pods, add custom checks (e.g., health endpoint status, request latency) using additional
UserParameterentries.
内容的提问来源于stack exchange,提问作者manu thankachan

