You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Lambda后端自定义资源CloudFormation模板CREATE_FAILED问题解决

解决CloudFormation自定义资源超时问题

你的堆栈报错“自定义资源未在预期时间内稳定”,核心原因是你的Lambda函数没有向CloudFormation发送自定义资源的响应信号。CloudFormation调用自定义资源后,会等待Lambda通过预签名URL返回成功/失败的回调,若超时未收到响应,就会判定堆栈失败。

关键修复步骤

  • 必须处理CloudFormation传递的ResponseURL,在Lambda执行完成(无论成功/失败)后发送响应
  • 添加异常捕获逻辑,确保出错时也能返回失败响应,避免CloudFormation无限等待
  • 优化Lambda执行逻辑,保证最终只会发送一次响应(即使处理多个目录)

修正后的完整CloudFormation模板

AWSTemplateFormatVersion: '2010-09-09'
Description: Lambda function to register event topic with existing directory ID
Parameters:
  RoleName:
    Type: String
    Description: "IAM Role used for Lambda execution"
    Default: "arn:aws:iam::<<Accountnumber>>:role/LambdaExecutionRole"
  EnvVariable:
    Type: String
    Description: "The Environment variable set for the lambda func"
    Default: "ESdirsvcSNS"
Resources:
  REGISTEREVENTTOPIC:
    Type: 'AWS::Lambda::Function'
    Properties:
      FunctionName: dirsvc_snstopic_lambda
      Handler: index.lambda_handler
      Runtime: python3.6
      Description: Lambda func code to assoc dirID with created SNS topic
      Code:
        ZipFile: |
          import boto3
          import os
          import logging
          import json
          import urllib3

          dsclient = boto3.client('ds')
          http = urllib3.PoolManager()

          def send_response(event, context, response_status, response_data):
              # 构造CloudFormation响应内容
              response_body = json.dumps({
                  "Status": response_status,
                  "Reason": f"See the details in CloudWatch Log Stream: {context.log_stream_name}",
                  "PhysicalResourceId": context.log_stream_name,
                  "StackId": event["StackId"],
                  "RequestId": event["RequestId"],
                  "LogicalResourceId": event["LogicalResourceId"],
                  "Data": response_data
              })

              # 发送响应到CloudFormation的预签名URL
              headers = {'Content-Type': ''}
              try:
                  response = http.request('PUT', event['ResponseURL'], body=response_body, headers=headers)
                  print(f"Response sent successfully, status code: {response.status}")
              except Exception as e:
                  print(f"Failed to send response: {str(e)}")

          def lambda_handler(event, context):
              response_data = {}
              try:
                  response = dsclient.describe_directories()
                  print(f"Directories found: {len(response['DirectoryDescriptions'])}")
                  registered_count = 0
                  for directory in response['DirectoryDescriptions']:
                      dir_id = directory['DirectoryId']
                      list_topics = dsclient.describe_event_topics(DirectoryId=dir_id)
                      event_topics = list_topics['EventTopics']
                      if len(event_topics) == 0:
                          dsclient.register_event_topic(
                              DirectoryId=dir_id,
                              TopicName=os.environ['MONITORING_TOPIC_NAME']
                          )
                          registered_count += 1
                          print(f"Registered topic for directory {dir_id}")
                      else:
                          print(f"Directory {dir_id} already has event topics configured")
                  response_data["RegisteredDirectories"] = registered_count
                  # 发送成功响应
                  send_response(event, context, "SUCCESS", response_data)
              except Exception as e:
                  logging.error(f"Error occurred: {str(e)}")
                  response_data["Error"] = str(e)
                  # 发送失败响应
                  send_response(event, context, "FAILED", response_data)
      Timeout: 60
      Environment:
        Variables:
          MONITORING_TOPIC_NAME: !Ref EnvVariable
      Role: !Ref RoleName
  InvokeLambda:
    Type: Custom::InvokeLambda
    Properties:
      ServiceToken: !GetAtt REGISTEREVENTTOPIC.Arn
      ReservedConcurrentExecutions: 1

修复细节说明

  1. 新增send_response函数:专门处理向CloudFormation发送响应的逻辑,包含必填的响应字段(Status、RequestId、StackId等)
  2. 异常捕获:用try-except包裹所有业务逻辑,出错时记录日志并发送FAILED响应
  3. 明确响应时机:无论处理多少个目录,最终都会发送一次成功/失败响应,确保CloudFormation能及时收到信号
  4. 使用urllib3发送请求:避免依赖requests库(Python 3.6默认不包含),保证Lambda能正常运行

另外,建议检查Lambda执行角色的权限:确保角色拥有ds:DescribeDirectories、ds:DescribeEventTopics、ds:RegisterEventTopic权限,以及CloudWatch Logs的写入权限(用于排查问题)。

内容的提问来源于stack exchange,提问作者CMR H

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:52:30