You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过YAML CloudFormation栈部署时自动上传ETL脚本至S3

问题

我一直在用YAML编写CloudFormation栈并部署到AWS基础设施(因遗留原因无法切换到CDK)。下面的YAML代码是CloudFormation栈的一部分,用于创建Glue Job,它通过ScriptLocation从S3存储桶加载transform_json_to_parquet.py脚本。当前方案的主要限制是必须确保transform_json_to_parquet.py脚本已经存在于<S3-bucket-name-1>中,所以我得手动把脚本上传到这个桶里。我想知道有没有办法在部署CloudFormation栈到AWS时自动加载这个脚本。

TransformJsonDataJob:
    Type: "AWS::Glue::Job"
    Properties:
      Role: !Ref AWSGlueETLJobRole  
      Name: "TransformJsonToParquet"
      Description: "Transform JSON to Parquet"
      Timeout: 5
      WorkerType: G.1X
      NumberOfWorkers: 2
      MaxRetries: 0
      Command:
        "Name": "glueetl"
        "ScriptLocation" : !Sub s3://<S3-bucket-name-1>/transform_json_to_parquet.py
      DefaultArguments: 
        "--s3_json_path" : !Sub s3://<S3-bucket-name-2>/
        "--s3_parquet_path" : !Sub s3://<S3-bucket-name-3>/
解决方案

下面是几种无需手动上传脚本,就能在部署CloudFormation栈时自动将脚本同步到目标S3桶的方法:

方法1:直接在CloudFormation中嵌入脚本内容

如果脚本内容不长,可以通过AWS::S3::Object资源将脚本代码直接写入模板,部署时自动创建S3对象,同时让Glue Job的ScriptLocation引用这个对象。

修改后的模板示例:

# 创建S3对象存储脚本
GlueScriptS3Object:
  Type: AWS::S3::Object
  Properties:
    Bucket: <S3-bucket-name-1> # 替换为你的目标桶名称
    Key: transform_json_to_parquet.py
    Content: |
      # 这里写入你的transform_json_to_parquet.py完整代码
      import sys
      from awsglue.transforms import *
      from awsglue.utils import getResolvedOptions
      from pyspark.context import SparkContext
      from awsglue.context import GlueContext
      from awsglue.job import Job

      args = getResolvedOptions(sys.argv, ['JOB_NAME', 's3_json_path', 's3_parquet_path'])
      sc = SparkContext()
      glueContext = GlueContext(sc)
      spark = glueContext.spark_session
      job = Job(glueContext)
      job.init(args['JOB_NAME'], args)

      # 读取JSON数据源
      dyf = glueContext.create_dynamic_frame.from_options(
          connection_type="s3",
          connection_options={"paths": [args['s3_json_path']]},
          format="json"
      )

      # 写入Parquet格式数据
      glueContext.write_dynamic_frame.from_options(
          frame=dyf,
          connection_type="s3",
          connection_options={"path": args['s3_parquet_path']},
          format="parquet"
      )

      job.commit()

# 修改Glue Job,引用自动创建的S3脚本
TransformJsonDataJob:
    Type: "AWS::Glue::Job"
    DependsOn: GlueScriptS3Object
    Properties:
      Role: !Ref AWSGlueETLJobRole  
      Name: "TransformJsonToParquet"
      Description: "Transform JSON to Parquet"
      Timeout: 5
      WorkerType: G.1X
      NumberOfWorkers: 2
      MaxRetries: 0
      Command:
        "Name": "glueetl"
        "ScriptLocation" : !Sub s3://<S3-bucket-name-1>/transform_json_to_parquet.py
      DefaultArguments: 
        "--s3_json_path" : !Sub s3://<S3-bucket-name-2>/
        "--s3_parquet_path" : !Sub s3://<S3-bucket-name-3>/

注意:如果脚本包含特殊字符(如引号)需要正确转义;脚本过长时会导致模板臃肿,不建议使用此方法。

方法2:用Lambda自定义资源自动上传脚本

如果脚本存放在本地或其他存储位置,可以通过Lambda自定义资源在栈部署阶段自动将脚本上传到目标S3桶。

模板示例片段:

# 自定义资源Lambda的执行角色
CustomResourceLambdaRole:
  Type: AWS::IAM::Role
  Properties:
    AssumeRolePolicyDocument:
      Version: '2012-10-17'
      Statement:
        - Effect: Allow
          Principal:
            Service: lambda.amazonaws.com
          Action: sts:AssumeRole
    ManagedPolicyArns:
      - arn:aws:iam::aws:policy/service-role/AWSLambdaBasicExecutionRole
    Policies:
      - PolicyName: S3UploadAccess
        PolicyDocument:
          Version: '2012-10-17'
          Statement:
            - Effect: Allow
              Action: s3:PutObject
              Resource: !Sub arn:aws:s3:::<S3-bucket-name-1>/*

# 负责上传脚本的Lambda函数
UploadGlueScriptLambda:
  Type: AWS::Lambda::Function
  Properties:
    Handler: index.lambda_handler
    Runtime: python3.11
    Role: !GetAtt CustomResourceLambdaRole.Arn
    Code:
      ZipFile: |
        import boto3
        import cfnresponse
        import os

        s3 = boto3.client('s3')
        bucket_name = os.environ['BUCKET_NAME']
        script_key = 'transform_json_to_parquet.py'

        def lambda_handler(event, context):
          try:
            if event['RequestType'] in ['Create', 'Update']:
              # 此处可替换为读取本地脚本、从其他S3桶拉取等逻辑
              script_content = """
              import sys
              from awsglue.transforms import *
              from awsglue.utils import getResolvedOptions
              from pyspark.context import SparkContext
              from awsglue.context import GlueContext
              from awsglue.job import Job

              args = getResolvedOptions(sys.argv, ['JOB_NAME', 's3_json_path', 's3_parquet_path'])
              sc = SparkContext()
              glueContext = GlueContext(sc)
              spark = glueContext.spark_session
              job = Job(glueContext)
              job.init(args['JOB_NAME'], args)

              dyf = glueContext.create_dynamic_frame.from_options(
                  connection_type="s3",
                  connection_options={"paths": [args['s3_json_path']]},
                  format="json"
              )

              glueContext.write_dynamic_frame.from_options(
                  frame=dyf,
                  connection_type="s3",
                  connection_options={"path": args['s3_parquet_path']},
                  format="parquet"
              )

              job.commit()
              """
              s3.put_object(Bucket=bucket_name, Key=script_key, Body=script_content.encode('utf-8'))
            cfnresponse.send(event, context, cfnresponse.SUCCESS, {})
          except Exception as e:
            cfnresponse.send(event, context, cfnresponse.FAILED, {'Error': str(e)})
    Environment:
      Variables:
        BUCKET_NAME: <S3-bucket-name-1>

# 触发Lambda上传脚本的自定义资源
UploadGlueScriptResource:
  Type: AWS::CloudFormation::CustomResource
  Properties:
    ServiceToken: !GetAtt UploadGlueScriptLambda.Arn

# 修改Glue Job,确保脚本上传完成后再创建
TransformJsonDataJob:
    Type: "AWS::Glue::Job"
    DependsOn: UploadGlueScriptResource
    Properties:
      Role: !Ref AWSGlueETLJobRole  
      Name: "TransformJsonToParquet"
      Description: "Transform JSON to Parquet"
      Timeout: 5
      WorkerType: G.1X
      NumberOfWorkers: 2
      MaxRetries: 0
      Command:
        "Name": "glueetl"
        "ScriptLocation" : !Sub s3://<S3-bucket-name-1>/transform_json_to_parquet.py
      DefaultArguments: 
        "--s3_json_path" : !Sub s3://<S3-bucket-name-2>/
        "--s3_parquet_path" : !Sub s3://<S3-bucket-name-3>/

方法3:用Shell脚本配合AWS CLI自动部署

如果不想修改CloudFormation模板,可以在部署栈前用脚本自动完成脚本上传和栈部署的流程:

# 先上传脚本到目标S3桶
aws s3 cp ./transform_json_to_parquet.py s3://<S3-bucket-name-1>/transform_json_to_parquet.py

# 然后部署CloudFormation栈
aws cloudformation deploy --template-file your-template.yaml --stack-name your-stack-name --capabilities CAPABILITY_IAM

这种方法简单直接,适合脚本频繁变动的场景,只需将上述命令放到同一个Shell脚本中执行即可。

内容的提问来源于stack exchange,提问作者Pankesh Patel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 11:34:50