使用Serverless部署时无法创建EFS的问题排查求助
问题描述
执行sls deploy --stage dev部署应用时,多数AWS资源已创建,但最终部署失败,报错如下:
✖ An error occurred: CheckDashindexDashsizeLambdaFunction - Resource handler returned message: "EFS file system arn:aws:elasticfilesystem:us-east-1:<account_id>:file-system/fs-<id1> referenced by access point arn:aws:elasticfilesystem:us-east-1:<account_id>:access-point/fsap-<id2> has mount targets created in all availability zones the function will execute in, but not all are in the available life cycle state yet. Please wait for them to become available and try the request again. (Service: Lambda, Status Code: 400, Request ID: ac1b6016-fd2d-4306-a7f1-745295b7cdb6)"
首次部署成功,但执行sls remove --stage dev清理资源后,6小时内重试10次部署均触发上述错误。相关环境变量已确认配置正确,serverless.yml配置如下:
org: ${env:ORG} service: lucene-serverless-${env:APP_NAME} variablesResolutionMode: 20210219 custom: name: ${sls:stage}-${self:service} region: ${opt:region, "us-east-1"} vpcId: ${env:LUCENE_SERVERLESS_VPC_ID} subnetId1: ${env:SUBNET_ID1} subnetId2: ${env:SUBNET_ID2} javaVersion: provided.al2 provider: name: aws profile: ${env:PROFILE} region: ${self:custom.region} versionFunctions: false apiGateway: shouldStartNameWithService: true tracing: lambda: false timeout: 15 environment: stage: prod DISABLE_SIGNAL_HANDLERS: true iam: role: statements: ${file(roleStatements.yml)} vpc: securityGroupIds: - Ref: EfsSecurityGroup subnetIds: - ${self:custom.subnetId1} - ${self:custom.subnetId2} package: individually: true functions: index: name: ${self:custom.name}-index runtime: ${self:custom.javaVersion} handler: native.handler reservedConcurrency: 1 memorySize: 256 timeout: 180 dependsOn: - EfsMountTarget1 - EfsMountTarget2 - EfsAccessPoint fileSystemConfig: localMountPath: /mnt/data arn: Fn::GetAtt: [EfsAccessPoint, Arn] package: artifact: target/function.zip environment: QUARKUS_LAMBDA_HANDLER: index QUARKUS_PROFILE: prod events: - sqs: arn: Fn::GetAtt: [WriteQueue, Arn] batchSize: 5000 maximumBatchingWindow: 5 enqueue-index: name: ${self:custom.name}-enqueue-index runtime: ${self:custom.javaVersion} handler: native.handler memorySize: 256 package: artifact: target/function.zip vpc: securityGroupIds: [] subnetIds: [] events: - http: POST /index environment: QUARKUS_LAMBDA_HANDLER: enqueue-index QUARKUS_PROFILE: prod QUEUE_URL: Ref: WriteQueue resources: Resources: WriteQueue: Type: AWS::SQS::Queue Properties: QueueName: ${self:custom.name}-write-queue VisibilityTimeout: 900 RedrivePolicy: deadLetterTargetArn: Fn::GetAtt: [WriteDLQ, Arn] maxReceiveCount: 5 WriteDLQ: Type: AWS::SQS::Queue Properties: QueueName: ${self:custom.name}-write-dlq MessageRetentionPeriod: 1209600 # 14 days in seconds FileSystem: Type: AWS::EFS::FileSystem Properties: BackupPolicy: Status: DISABLED FileSystemTags: - Key: Name Value: ${self:custom.name}-fs PerformanceMode: generalPurpose ThroughputMode: elastic # faster scale up/down Encrypted: true FileSystemPolicy: Version: "2012-10-17" Statement: - Effect: "Allow" Action: - "elasticfilesystem:ClientMount" Principal: AWS: "*" EfsSecurityGroup: Type: AWS::EC2::SecurityGroup Properties: VpcId: ${self:custom.vpcId} GroupDescription: "mnt target sg" SecurityGroupIngress: - IpProtocol: -1 CidrIp: "0.0.0.0/0" - IpProtocol: -1 CidrIpv6: "::/0" SecurityGroupEgress: - IpProtocol: -1 CidrIp: "0.0.0.0/0" - IpProtocol: -1 CidrIpv6: "::/0" EfsMountTarget1: Type: AWS::EFS::MountTarget Properties: FileSystemId: !Ref FileSystem SubnetId: ${self:custom.subnetId1} SecurityGroups: - Ref: EfsSecurityGroup EfsMountTarget2: Type: AWS::EFS::MountTarget Properties: FileSystemId: !Ref FileSystem SubnetId: ${self:custom.subnetId2} SecurityGroups: - Ref: EfsSecurityGroup EfsAccessPoint: Type: "AWS::EFS::AccessPoint" Properties: FileSystemId: !Ref FileSystem PosixUser: Uid: "1000" Gid: "1000" RootDirectory: CreationInfo: OwnerGid: "1000" OwnerUid: "1000" Permissions: "0777" Path: "/mnt/data"
问题分析与解决方案
问题原因
报错核心是Lambda创建时,关联的EFS挂载目标虽已完成资源创建,但未进入available生命周期状态。CloudFormation仅确认AWS::EFS::MountTarget资源创建动作完成,不会等待其实际状态变为可用,导致Lambda尝试挂载时触发错误。首次部署成功是因为资源创建节奏刚好让挂载目标在Lambda创建前就绪,但清理后重新部署时,CloudFormation创建挂载目标的速度快于其实际就绪速度。
修复方案
添加自定义资源,强制等待所有EFS挂载目标进入available状态后,再创建Lambda函数。
步骤1:在serverless.yml的resources中添加自定义检查资源
resources: Resources: # ... 保留原有所有资源,新增以下内容 ... CheckMountTargetsAvailable: Type: Custom::CheckMountTargets Properties: ServiceToken: !GetAtt CheckMountTargetsFunction.Arn FileSystemId: !Ref FileSystem SubnetIds: - ${self:custom.subnetId1} - ${self:custom.subnetId2} CheckMountTargetsFunction: Type: AWS::Lambda::Function Properties: Runtime: python3.11 Handler: index.lambda_handler Role: !GetAtt CheckMountTargetsRole.Arn Code: ZipFile: | import boto3 import time import cfnresponse efs = boto3.client('efs') def lambda_handler(event, context): try: file_system_id = event['ResourceProperties']['FileSystemId'] subnet_ids = event['ResourceProperties']['SubnetIds'] # 循环检查挂载目标状态,直到全部可用 while True: response = efs.describe_mount_targets(FileSystemId=file_system_id) available_targets = [ target for target in response['MountTargets'] if target['SubnetId'] in subnet_ids and target['LifeCycleState'] == 'available' ] if len(available_targets) == len(subnet_ids): break time.sleep(10) cfnresponse.send(event, context, cfnresponse.SUCCESS, {}) except Exception as e: cfnresponse.send(event, context, cfnresponse.FAILED, {'Error': str(e)}) CheckMountTargetsRole: Type: AWS::IAM::Role Properties: AssumeRolePolicyDocument: Version: '2012-10-17' Statement: - Effect: Allow Principal: Service: lambda.amazonaws.com Action: sts:AssumeRole Policies: - PolicyName: CheckMountTargetsPolicy PolicyDocument: Version: '2012-10-17' Statement: - Effect: Allow Action: - efs:DescribeMountTargets Resource: '*' - Effect: Allow Action: - logs:CreateLogGroup - logs:CreateLogStream - logs:PutLogEvents Resource: arn:aws:logs:*:*:*
步骤2:修改index函数的dependsOn配置
将原有依赖挂载目标的配置,替换为依赖自定义检查资源:
functions: index: # ... 保留其他所有配置 ... dependsOn: - CheckMountTargetsAvailable - EfsAccessPoint
可选验证步骤
若重试仍失败,可手动登录AWS控制台,检查EFS挂载目标的生命周期状态,若存在卡在creating或deleting状态的目标,手动清理后重新部署。
内容的提问来源于stack exchange,提问作者Cerin
相关产品推荐
相关产品推荐

