You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Prometheus嵌套查询需求:集群指标超限后定位主机并触发脚本

How to Implement Nested Queries in Prometheus + Trigger Host-Specific Scripts

Got it, let's break down exactly how to set this up—from crafting targeted PromQL queries to configuring alerts that run your specified script when cluster thresholds are breached and specific hosts are the culprits.

1. Build Cluster + Host-Level PromQL Queries

First, you need two layers of queries: one to detect when the cluster-wide metric crosses your threshold, and another to pinpoint which individual hosts are driving that breach.

Example Scenario (CPU Usage)

Let's say we want to alert when the cluster's average CPU usage exceeds 80%, then identify all hosts where their individual CPU usage is over 75%:

  • Cluster-level check (triggers the investigation):

    avg by (cluster) (100 - (avg by (instance, cluster) (irate(node_cpu_seconds_total{mode="idle"}[1m])) * 100)) > 80
    

    This calculates the average CPU usage across all hosts in the cluster and flags if it’s over 80%.

  • Host-level check (finds problematic hosts):

    100 - (avg by (instance) (irate(node_cpu_seconds_total{mode="idle"}[1m])) * 100) > 75
    

    This targets individual instances (hosts) where CPU usage exceeds 75%.

To tie these together, we’ll embed the host-level query directly into our alert rule so Alertmanager gets the list of affected hosts when the cluster alert fires.

2. Configure Prometheus Alert Rules

Create an alert rule file (e.g., cluster_alerts.yml) that combines both checks and includes host details in the alert payload:

groups:
- name: cluster_resource_alerts
  rules:
  - alert: ClusterHighCPUUsage
    expr: avg by (cluster) (100 - (avg by (instance, cluster) (irate(node_cpu_seconds_total{mode="idle"}[1m])) * 100)) > 80
    for: 2m  # Wait 2 minutes to avoid false positives
    labels:
      severity: critical
    annotations:
      summary: "Cluster {{ $labels.cluster }} has high CPU usage"
      description: "Cluster average CPU is {{ $value | round 2 }}%. Problematic hosts: {{ range query '100 - (avg by (instance) (irate(node_cpu_seconds_total{mode=\"idle\"}[1m])) * 100) > 75' }}{{ .Labels.instance }} ({{ .Value | round 2 }}%), {{ end }}"

The query function in annotations pulls the host-level results directly into the alert, so you’ll have a full list of affected hosts when the alert triggers.

3. Set Up Alertmanager to Run Your Script

Now configure Alertmanager to execute your specified script when the alert fires. Here are two reliable approaches:

Build a simple web service (e.g., Python Flask, Go) that listens for Alertmanager webhook requests, parses the problematic hosts from the alert annotations, and runs your script for each host.

Alertmanager Config Snippet:

route:
  group_by: ['cluster']
  receiver: 'host-script-executor'

receivers:
- name: 'host-script-executor'
  webhook_configs:
  - url: 'http://your-webhook-service:5000/run-target-script'
    send_resolved: false

Your webhook service can extract hosts from the description annotation and run commands like:

./your-specified-script.sh {{ host_name_or_ip }}

Option 2: Local Executable Receiver

If you prefer running scripts directly on the Alertmanager server, use the exec receiver (requires enabling the feature flag in newer Alertmanager versions):

Alertmanager Config Snippet:

receivers:
- name: 'host-script-executor'
  exec_configs:
  - command: '/path/to/your-script-wrapper.sh'
    args: ['{{ .CommonAnnotations.description }}']

Your wrapper script can parse the host list from the input argument and execute the target script for each affected host.

Key Tips for Reliability

  • Test queries first: Validate both cluster and host-level PromQL queries in the Prometheus UI before deploying alerts.
  • Add deduplication: Prevent repeated script runs for the same host by using lock files or tracking execution times in a temp database.
  • Log everything: Add logging to your script/webhook to debug failed executions.

内容的提问来源于stack exchange,提问作者Pranjal Gupta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:55:07