CloudWatch指标告警抑制:如何避免阈值触发后重复发送相同错误通知
Got it, I’ve dealt with this exact scenario before—nothing’s more annoying than getting spammed with the same alert over and over while you’re fixing the issue. Here are three solid approaches to implement "one-time notification" for your pattern-matched CloudWatch alarms:
1. Use CloudWatch Alarm’s Built-in Action Suppression
The easiest way is to leverage AWS’s native suppression feature, which lets you stop repeat notifications after the first alert is triggered. Here’s how to set it up:
- Open your CloudWatch Alarm in the AWS Console.
- Go to the Actions tab, then scroll down to Suppress actions.
- Enable the toggle, then set a Suppression period (e.g., 24 hours, or however long you want to avoid repeats for the same alarm state).
- Choose whether to suppress actions only when the alarm remains in the same state (this is the default, perfect for your use case—once it triggers ALARM, no more notifications until it goes back to OK and re-triggers).
Note: This works at the alarm level. If your pattern-matched alarm covers multiple resources (e.g., all EC2 instances with a specific tag), this will suppress notifications for the entire alarm, not individual resource errors. If you need per-resource or per-error granularity, skip to the next method.
2. Build Custom Deduplication with SNS + Lambda + DynamoDB
For more control (like suppressing based on specific error details instead of just alarm state), build a custom pipeline:
Step 1: Redirect Alerts to an SNS Topic
- Update your CloudWatch Alarm to send notifications to an intermediate SNS topic (instead of sending directly to your email/Slack).
Step 2: Create a Deduplication Database
- Spin up a DynamoDB table with a primary key (e.g.,
error_identifier) and aTTLattribute to auto-expire old entries. Theerror_identifiershould be a unique string for each distinct error—like combining the alarm name, resource ID, and error code from the alert payload.
Step 3: Write a Lambda Function to Handle Deduplication
- Subscribe the Lambda function to your SNS topic. The function logic should:
- Parse the incoming CloudWatch alert payload to extract key details (resource ID, error message, alarm name).
- Generate the
error_identifierhash/string. - Query DynamoDB to check if this identifier already exists.
- If it doesn’t exist:
- Forward the alert to your actual notification channel (another SNS topic for email/Slack, or direct API call).
- Write the
error_identifierto DynamoDB with a TTL (e.g., 86400 seconds for 24 hours).
- If it does exist: Skip sending the notification.
Pro Tip: Add logging to the Lambda function so you can track which alerts were suppressed and why.
3. Combine Alarm States with State Updates
If you want notifications to resume only after the issue is resolved, you can pair the built-in suppression with alarm state transitions:
- Set the suppression period to a long value (e.g., 7 days), but configure the alarm to reset the suppression when it transitions back to
OK. - This way, if the same error reoccurs after the issue was fixed, you’ll get a fresh notification—but while it’s stuck in
ALARM, no repeats.
Each approach has its tradeoffs: the built-in suppression is fastest to set up, while the Lambda pipeline gives you full control over what counts as a "duplicate" error. Pick the one that fits your use case!
内容的提问来源于stack exchange,提问作者Unicorn

