You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于CloudWatch Insights查询结果触发CloudWatch告警?EC2 CPU利用率超阈值场景的告警实现方案

Hey there! Let's tackle your two CloudWatch questions with practical, actionable steps tailored to your needs.

1. How to Trigger CloudWatch Alarms Using CloudWatch Insights Queries

First off, it’s important to note that CloudWatch Insights (a log/metric query tool) doesn’t directly trigger alarms out of the box. To link Insights queries to alarm actions, you’ll need to use AWS services like EventBridge and Lambda to bridge the gap. Here’s a step-by-step workflow:

  • Schedule your Insights query: Create an EventBridge rule that runs on a recurring schedule (e.g., every 5 minutes). This rule will initiate the execution of your Insights query on a regular basis.
  • Route results to Lambda: Set the target of your EventBridge rule to a Lambda function. This function will handle running the query, fetching results, and triggering alarm actions.
  • Process results and trigger alarms: Inside your Lambda function:
    1. Use the CloudWatch API (start_query and get_query_results) to run your Insights query and retrieve the output.
    2. Validate if the results meet your alarm conditions (e.g., non-empty results, specific metric values).
    3. Choose an alarm action based on your needs:
      • Persistent monitoring: Use the put_metric_alarm API to create or update a CloudWatch Alarm tied to the relevant resource (like an EC2 instance). This is ideal for ongoing threshold monitoring.
      • Immediate alerts: Use SNS to send direct notifications (email, Slack, etc.) without creating a persistent alarm—great for one-time or ad-hoc alerting.

A simplified Python snippet for the Lambda function:

import boto3
import time

cloudwatch = boto3.client('cloudwatch')
sns = boto3.client('sns')

def lambda_handler(event, context):
    # Start the Insights query
    query_response = cloudwatch.start_query(
        logGroupName='/aws/ec2/your-log-group',
        queryString='YOUR_INSIGHTS_QUERY_HERE',
        startTime=int(time.time()) - 300,  # Last 5 minutes
        endTime=int(time.time())
    )
    query_id = query_response['queryId']
    
    # Wait for query completion (add retry logic for production use)
    time.sleep(5)
    results = cloudwatch.get_query_results(queryId=query_id)
    
    # Trigger alert if results exist
    if len(results['results']) > 1:  # Skip header row
        sns.publish(
            TopicArn='arn:aws:sns:us-east-1:123456789012:your-alert-topic',
            Message=f"Insights query triggered an alert: {results['results'][1:]}"
        )
2. Optimal Way to Trigger Alarms for Each Record from Your EC2 CPU Query

Your query SELECT CPUUtilization FROM SCHEMA("AWS/EC2", InstanceId) WHERE CPUUtilization > 80 GROUP BY InstanceId targets EC2 instances with high CPU usage. Let’s break down the best approaches based on your use case:

Option 1: Native CloudWatch Metric Alarms (Simplest & Most Efficient)

Since you’re working with EC2’s built-in CPUUtilization metric, the optimal solution is to skip Insights entirely and use CloudWatch’s native metric alarms. This avoids unnecessary complexity and leverages AWS’s built-in monitoring:

  • Create a dimensioned alarm:
    1. Navigate to CloudWatch Alarms > Create alarm.
    2. Select the AWS/EC2 namespace, then the CPUUtilization metric.
    3. Choose the InstanceId dimension, and select "All instances" (or specific ones) to apply the alarm to every instance.
    4. Set the threshold to > 80, configure your evaluation period (e.g., 5 consecutive minutes), and link to an SNS topic for notifications.

This approach automatically triggers an alarm for each individual instance that exceeds the CPU threshold—no extra services required.

Option 2: Insights + Lambda (For Complex Query Scenarios)

If you need this logic as part of a more complex Insights workflow (e.g., combining metrics with log data), use the EventBridge + Lambda setup from the first question, with these tweaks:

  • Process each record individually:
    Loop through each result in the Insights output, extract the InstanceId and CPUUtilization value, then trigger an action for each instance.
    def process_insights_results(results):
        # Skip the header row
        for record in results['results'][1:]:
            instance_id = next(item['value'] for item in record if item['field'] == 'InstanceId')
            cpu_util = next(item['value'] for item in record if item['field'] == 'CPUUtilization')
            
            # Send direct notification for the instance
            sns.publish(
                TopicArn='arn:aws:sns:us-east-1:123456789012:your-alert-topic',
                Message=f"ALERT: Instance {instance_id} has CPU utilization at {cpu_util}%"
            )
            
            # Or create a persistent alarm for ongoing monitoring
            cloudwatch.put_metric_alarm(
                AlarmName=f"High-CPU-Alarm-{instance_id}",
                AlarmDescription=f"CPU utilization exceeds 80% for {instance_id}",
                ActionsEnabled=True,
                AlarmActions=['arn:aws:sns:us-east-1:123456789012:your-alert-topic'],
                MetricName='CPUUtilization',
                Namespace='AWS/EC2',
                Statistic='Average',
                Dimensions=[{'Name': 'InstanceId', 'Value': instance_id}],
                Period=300,
                EvaluationPeriods=1,
                Threshold=80.0,
                ComparisonOperator='GreaterThanThreshold'
            )
    
  • Add deduplication: To avoid spamming notifications, store triggered instance IDs in DynamoDB with a timestamp, and skip alerts if the instance was already notified within a cooling period (e.g., 15 minutes).

内容的提问来源于stack exchange,提问作者Biju

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 19:09:04