Prometheus嵌套查询需求:集群指标超限后定位主机并触发脚本
Got it, let's break down exactly how to set this up—from crafting targeted PromQL queries to configuring alerts that run your specified script when cluster thresholds are breached and specific hosts are the culprits.
1. Build Cluster + Host-Level PromQL Queries
First, you need two layers of queries: one to detect when the cluster-wide metric crosses your threshold, and another to pinpoint which individual hosts are driving that breach.
Example Scenario (CPU Usage)
Let's say we want to alert when the cluster's average CPU usage exceeds 80%, then identify all hosts where their individual CPU usage is over 75%:
Cluster-level check (triggers the investigation):
avg by (cluster) (100 - (avg by (instance, cluster) (irate(node_cpu_seconds_total{mode="idle"}[1m])) * 100)) > 80This calculates the average CPU usage across all hosts in the cluster and flags if it’s over 80%.
Host-level check (finds problematic hosts):
100 - (avg by (instance) (irate(node_cpu_seconds_total{mode="idle"}[1m])) * 100) > 75This targets individual instances (hosts) where CPU usage exceeds 75%.
To tie these together, we’ll embed the host-level query directly into our alert rule so Alertmanager gets the list of affected hosts when the cluster alert fires.
2. Configure Prometheus Alert Rules
Create an alert rule file (e.g., cluster_alerts.yml) that combines both checks and includes host details in the alert payload:
groups: - name: cluster_resource_alerts rules: - alert: ClusterHighCPUUsage expr: avg by (cluster) (100 - (avg by (instance, cluster) (irate(node_cpu_seconds_total{mode="idle"}[1m])) * 100)) > 80 for: 2m # Wait 2 minutes to avoid false positives labels: severity: critical annotations: summary: "Cluster {{ $labels.cluster }} has high CPU usage" description: "Cluster average CPU is {{ $value | round 2 }}%. Problematic hosts: {{ range query '100 - (avg by (instance) (irate(node_cpu_seconds_total{mode=\"idle\"}[1m])) * 100) > 75' }}{{ .Labels.instance }} ({{ .Value | round 2 }}%), {{ end }}"
The query function in annotations pulls the host-level results directly into the alert, so you’ll have a full list of affected hosts when the alert triggers.
3. Set Up Alertmanager to Run Your Script
Now configure Alertmanager to execute your specified script when the alert fires. Here are two reliable approaches:
Option 1: Webhook Receiver (Recommended)
Build a simple web service (e.g., Python Flask, Go) that listens for Alertmanager webhook requests, parses the problematic hosts from the alert annotations, and runs your script for each host.
Alertmanager Config Snippet:
route: group_by: ['cluster'] receiver: 'host-script-executor' receivers: - name: 'host-script-executor' webhook_configs: - url: 'http://your-webhook-service:5000/run-target-script' send_resolved: false
Your webhook service can extract hosts from the description annotation and run commands like:
./your-specified-script.sh {{ host_name_or_ip }}
Option 2: Local Executable Receiver
If you prefer running scripts directly on the Alertmanager server, use the exec receiver (requires enabling the feature flag in newer Alertmanager versions):
Alertmanager Config Snippet:
receivers: - name: 'host-script-executor' exec_configs: - command: '/path/to/your-script-wrapper.sh' args: ['{{ .CommonAnnotations.description }}']
Your wrapper script can parse the host list from the input argument and execute the target script for each affected host.
Key Tips for Reliability
- Test queries first: Validate both cluster and host-level PromQL queries in the Prometheus UI before deploying alerts.
- Add deduplication: Prevent repeated script runs for the same host by using lock files or tracking execution times in a temp database.
- Log everything: Add logging to your script/webhook to debug failed executions.
内容的提问来源于stack exchange,提问作者Pranjal Gupta

