You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何创建涉及多指标的AWS CloudWatch告警:API网关与Lambda场景

Setting Up CloudWatch Alarms for API Gateway & Lambda

Let’s break down exactly how to configure these CloudWatch Alarms using Metric Math—this is a common setup I’ve implemented for production APIs, so I’ll walk you through each step clearly.

1. API Gateway Alarm: 5XX Error Rate ≥10% (5-Minute Window)

To trigger an alert when 10% or more of requests result in 5XX errors over 5 minutes, you’ll need to calculate the error rate using two base metrics combined with a math expression:

  • First, add your API Gateway’s 5XXError metric, set the statistic to Sum, and choose a 5-minute period.
  • Next, add the Count metric (total requests) for the same API/stage, also set to Sum with a 5-minute period.
  • Create a Metric Math expression like this:
    ERROR_RATE = m1 / m2
    
    Where m1 is the reference ID for your 5XXError Sum metric, and m2 is the ID for your total Count Sum metric.
  • Set your alarm condition to trigger when ERROR_RATE is greater than or equal to 0.1 (10%).
  • Finish by configuring your notification target (e.g., an SNS topic for email/Slack alerts) and other standard alarm settings.

2. API Gateway Alarm: 6% of Requests Exceed 4-Second Latency

For this alert, we need to track both the percentage of slow requests and confirm the maximum latency threshold is hit. Here’s how:

  • Add a Latency metric for your API/stage, set the statistic to Count, and add a filter: Latency > 4000 (since latency is tracked in milliseconds). Use a 5-minute period. This counts how many requests took longer than 4 seconds.
  • Add the total Count metric again (same as step 1) to get overall request volume.
  • Create a Metric Math expression for the slow request rate:
    SLOW_REQUEST_RATE = m1 / m2
    
    Where m1 is the count of requests with latency >4000, and m2 is total requests.
  • Optionally, add a second Latency metric with statistic Maximum to confirm the peak latency exceeds 4000ms.
  • Set your alarm condition to trigger when both SLOW_REQUEST_RATE ≥ 0.06 (6%) and MAX_LATENCY > 4000 are true. This ensures you’re only alerted when a meaningful portion of requests are slow, not just a single outlier.

3. Quick Lambda Alarm Configurations (Since You Mentioned It)

While your main ask is API Gateway, here are a couple of critical Lambda alarms you might want to set up alongside:

  • Error Rate Alarm: Calculate Errors Sum / Invocations Sum, trigger when the rate exceeds your threshold (e.g., 5%).
  • Throttle Alarm: Track Throttles Sum / Invocations Sum—throttles are a sign you need to adjust concurrency limits.
  • Duration Alarm: Monitor the P95 or Maximum of the Duration metric, and alert if it’s approaching your function’s timeout (e.g., 80% of the timeout value).

内容的提问来源于stack exchange,提问作者Darshan Ambhaikar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:13:59