GCP Stackdriver Pub/Sub告警配置问题:AND_WITH_MATCHING_RESOURCE不符合预期
Got it, let's break down how to fix your Stackdriver alerting setup—since you're coming from a Prometheus background, I'll draw direct parallels to help you map your familiar logic to Stackdriver's model.
First, Why Your Current Setup Is Triggering False Positives
Your choice of AND_WITH_MATCHING_RESOURCE is the right call (it's Stackdriver's equivalent of Prometheus's label-based grouping to match time series across metrics), but a few misconfigurations are causing unintended alerts:
- The ack rate threshold of
1is way too high for your "接近0" (near zero) requirement - Using
REDUCE_SUMfor cross-series aggregation might be combining metrics across multiple subscriptions, masking individual subscription behavior - The alignment and grouping rules weren't fully consistent between the two conditions, leading to mismatched resource evaluations
Mapping Prometheus Logic to Stackdriver
In Prometheus, you'd write something like this to target the exact scenario you want:
# Low consumption rate (near 0 ack messages per second) sum by(topic_id, subscription_id) (rate(pubsub_subscription_ack_message_count[1m])) < 0.1 # AND high unacked messages sum by(topic_id, subscription_id) (pubsub_subscription_num_undelivered_messages) > XXX
Here's how to translate that to Stackdriver correctly:
Corrected Alert Configuration
{ "combiner": "AND_WITH_MATCHING_RESOURCE", "conditions": [ { "conditionThreshold": { "aggregations": [ { "alignmentPeriod": "60s", "crossSeriesReducer": "REDUCE_NONE", "groupByFields": [ "metadata.system_labels.topic_id", "resource.label.subscription_id" ], "perSeriesAligner": "ALIGN_RATE" } ], "comparison": "COMPARISON_LT", "duration": "300s", "filter": "metric.type=\"pubsub.googleapis.com/subscription/ack_message_count\" resource.type=\"pubsub_subscription\" resource.label.\"project_id\"=\"pl-service-prod-lm-fr\"", "thresholdValue": 0.1, "trigger": { "count": 1 } }, "displayName": "Ack message rate (near zero)" }, { "conditionThreshold": { "aggregations": [ { "alignmentPeriod": "60s", "crossSeriesReducer": "REDUCE_NONE", "groupByFields": [ "metadata.system_labels.topic_id", "resource.label.subscription_id" ], "perSeriesAligner": "ALIGN_LATEST" } ], "comparison": "COMPARISON_GT", "duration": "300s", "filter": "metric.type=\"pubsub.googleapis.com/subscription/num_undelivered_messages\" resource.type=\"pubsub_subscription\" resource.label.\"project_id\"=\"pl-service-prod-lm-fr\"", "thresholdValue": 100, // Replace with your XXX value "trigger": { "count": 1 } }, "displayName": "Unacked messages exceed threshold" } ], "displayName": "Pub/Sub Subscription Stalled (High Unacked + Low Consumption)", "enabled": true, "incidentStrategy": {} }
Key Changes Explained
- Threshold Tweak: Lowered the ack rate threshold to
0.1(adjust this to your exact definition of "接近0"—e.g.,0.05for even lower activity) - Cross-Series Reducer: Switched from
REDUCE_SUMtoREDUCE_NONEto evaluate each subscription individually, matching Prometheus's per-label grouping behavior - Consistent Grouping: Both conditions use identical
groupByFieldsto ensure Stackdriver matches the exact same subscription/topic pairs across both metrics - Unacked Message Aligner: Used
ALIGN_LATESTfornum_undelivered_messagessince we care about the current number of unacked messages, not an average over time - Clear Naming: Updated display names to make each condition's purpose immediately obvious
Verifying the Logic
AND_WITH_MATCHING_RESOURCE will now only trigger an alert if both conditions are true for the same subscription/topic pair for the full 300-second duration—exactly what you need: high unacked messages, and almost no consumption activity.
内容的提问来源于stack exchange,提问作者Maxime Wojtczak

