Lambda后端自定义资源CloudFormation模板CREATE_FAILED问题解决
解决CloudFormation自定义资源超时问题
你的堆栈报错“自定义资源未在预期时间内稳定”,核心原因是你的Lambda函数没有向CloudFormation发送自定义资源的响应信号。CloudFormation调用自定义资源后,会等待Lambda通过预签名URL返回成功/失败的回调,若超时未收到响应,就会判定堆栈失败。
关键修复步骤
- 必须处理CloudFormation传递的
ResponseURL,在Lambda执行完成(无论成功/失败)后发送响应 - 添加异常捕获逻辑,确保出错时也能返回失败响应,避免CloudFormation无限等待
- 优化Lambda执行逻辑,保证最终只会发送一次响应(即使处理多个目录)
修正后的完整CloudFormation模板
AWSTemplateFormatVersion: '2010-09-09' Description: Lambda function to register event topic with existing directory ID Parameters: RoleName: Type: String Description: "IAM Role used for Lambda execution" Default: "arn:aws:iam::<<Accountnumber>>:role/LambdaExecutionRole" EnvVariable: Type: String Description: "The Environment variable set for the lambda func" Default: "ESdirsvcSNS" Resources: REGISTEREVENTTOPIC: Type: 'AWS::Lambda::Function' Properties: FunctionName: dirsvc_snstopic_lambda Handler: index.lambda_handler Runtime: python3.6 Description: Lambda func code to assoc dirID with created SNS topic Code: ZipFile: | import boto3 import os import logging import json import urllib3 dsclient = boto3.client('ds') http = urllib3.PoolManager() def send_response(event, context, response_status, response_data): # 构造CloudFormation响应内容 response_body = json.dumps({ "Status": response_status, "Reason": f"See the details in CloudWatch Log Stream: {context.log_stream_name}", "PhysicalResourceId": context.log_stream_name, "StackId": event["StackId"], "RequestId": event["RequestId"], "LogicalResourceId": event["LogicalResourceId"], "Data": response_data }) # 发送响应到CloudFormation的预签名URL headers = {'Content-Type': ''} try: response = http.request('PUT', event['ResponseURL'], body=response_body, headers=headers) print(f"Response sent successfully, status code: {response.status}") except Exception as e: print(f"Failed to send response: {str(e)}") def lambda_handler(event, context): response_data = {} try: response = dsclient.describe_directories() print(f"Directories found: {len(response['DirectoryDescriptions'])}") registered_count = 0 for directory in response['DirectoryDescriptions']: dir_id = directory['DirectoryId'] list_topics = dsclient.describe_event_topics(DirectoryId=dir_id) event_topics = list_topics['EventTopics'] if len(event_topics) == 0: dsclient.register_event_topic( DirectoryId=dir_id, TopicName=os.environ['MONITORING_TOPIC_NAME'] ) registered_count += 1 print(f"Registered topic for directory {dir_id}") else: print(f"Directory {dir_id} already has event topics configured") response_data["RegisteredDirectories"] = registered_count # 发送成功响应 send_response(event, context, "SUCCESS", response_data) except Exception as e: logging.error(f"Error occurred: {str(e)}") response_data["Error"] = str(e) # 发送失败响应 send_response(event, context, "FAILED", response_data) Timeout: 60 Environment: Variables: MONITORING_TOPIC_NAME: !Ref EnvVariable Role: !Ref RoleName InvokeLambda: Type: Custom::InvokeLambda Properties: ServiceToken: !GetAtt REGISTEREVENTTOPIC.Arn ReservedConcurrentExecutions: 1
修复细节说明
- 新增
send_response函数:专门处理向CloudFormation发送响应的逻辑,包含必填的响应字段(Status、RequestId、StackId等) - 异常捕获:用
try-except包裹所有业务逻辑,出错时记录日志并发送FAILED响应 - 明确响应时机:无论处理多少个目录,最终都会发送一次成功/失败响应,确保CloudFormation能及时收到信号
- 使用
urllib3发送请求:避免依赖requests库(Python 3.6默认不包含),保证Lambda能正常运行
另外,建议检查Lambda执行角色的权限:确保角色拥有ds:DescribeDirectories、ds:DescribeEventTopics、ds:RegisterEventTopic权限,以及CloudWatch Logs的写入权限(用于排查问题)。
内容的提问来源于stack exchange,提问作者CMR H
相关产品推荐
相关产品推荐

