如何通过YAML CloudFormation栈部署时自动上传ETL脚本至S3
问题
我一直在用YAML编写CloudFormation栈并部署到AWS基础设施(因遗留原因无法切换到CDK)。下面的YAML代码是CloudFormation栈的一部分,用于创建Glue Job,它通过ScriptLocation从S3存储桶加载transform_json_to_parquet.py脚本。当前方案的主要限制是必须确保transform_json_to_parquet.py脚本已经存在于<S3-bucket-name-1>中,所以我得手动把脚本上传到这个桶里。我想知道有没有办法在部署CloudFormation栈到AWS时自动加载这个脚本。
TransformJsonDataJob: Type: "AWS::Glue::Job" Properties: Role: !Ref AWSGlueETLJobRole Name: "TransformJsonToParquet" Description: "Transform JSON to Parquet" Timeout: 5 WorkerType: G.1X NumberOfWorkers: 2 MaxRetries: 0 Command: "Name": "glueetl" "ScriptLocation" : !Sub s3://<S3-bucket-name-1>/transform_json_to_parquet.py DefaultArguments: "--s3_json_path" : !Sub s3://<S3-bucket-name-2>/ "--s3_parquet_path" : !Sub s3://<S3-bucket-name-3>/
解决方案
下面是几种无需手动上传脚本,就能在部署CloudFormation栈时自动将脚本同步到目标S3桶的方法:
方法1:直接在CloudFormation中嵌入脚本内容
如果脚本内容不长,可以通过AWS::S3::Object资源将脚本代码直接写入模板,部署时自动创建S3对象,同时让Glue Job的ScriptLocation引用这个对象。
修改后的模板示例:
# 创建S3对象存储脚本 GlueScriptS3Object: Type: AWS::S3::Object Properties: Bucket: <S3-bucket-name-1> # 替换为你的目标桶名称 Key: transform_json_to_parquet.py Content: | # 这里写入你的transform_json_to_parquet.py完整代码 import sys from awsglue.transforms import * from awsglue.utils import getResolvedOptions from pyspark.context import SparkContext from awsglue.context import GlueContext from awsglue.job import Job args = getResolvedOptions(sys.argv, ['JOB_NAME', 's3_json_path', 's3_parquet_path']) sc = SparkContext() glueContext = GlueContext(sc) spark = glueContext.spark_session job = Job(glueContext) job.init(args['JOB_NAME'], args) # 读取JSON数据源 dyf = glueContext.create_dynamic_frame.from_options( connection_type="s3", connection_options={"paths": [args['s3_json_path']]}, format="json" ) # 写入Parquet格式数据 glueContext.write_dynamic_frame.from_options( frame=dyf, connection_type="s3", connection_options={"path": args['s3_parquet_path']}, format="parquet" ) job.commit() # 修改Glue Job,引用自动创建的S3脚本 TransformJsonDataJob: Type: "AWS::Glue::Job" DependsOn: GlueScriptS3Object Properties: Role: !Ref AWSGlueETLJobRole Name: "TransformJsonToParquet" Description: "Transform JSON to Parquet" Timeout: 5 WorkerType: G.1X NumberOfWorkers: 2 MaxRetries: 0 Command: "Name": "glueetl" "ScriptLocation" : !Sub s3://<S3-bucket-name-1>/transform_json_to_parquet.py DefaultArguments: "--s3_json_path" : !Sub s3://<S3-bucket-name-2>/ "--s3_parquet_path" : !Sub s3://<S3-bucket-name-3>/
注意:如果脚本包含特殊字符(如引号)需要正确转义;脚本过长时会导致模板臃肿,不建议使用此方法。
方法2:用Lambda自定义资源自动上传脚本
如果脚本存放在本地或其他存储位置,可以通过Lambda自定义资源在栈部署阶段自动将脚本上传到目标S3桶。
模板示例片段:
# 自定义资源Lambda的执行角色 CustomResourceLambdaRole: Type: AWS::IAM::Role Properties: AssumeRolePolicyDocument: Version: '2012-10-17' Statement: - Effect: Allow Principal: Service: lambda.amazonaws.com Action: sts:AssumeRole ManagedPolicyArns: - arn:aws:iam::aws:policy/service-role/AWSLambdaBasicExecutionRole Policies: - PolicyName: S3UploadAccess PolicyDocument: Version: '2012-10-17' Statement: - Effect: Allow Action: s3:PutObject Resource: !Sub arn:aws:s3:::<S3-bucket-name-1>/* # 负责上传脚本的Lambda函数 UploadGlueScriptLambda: Type: AWS::Lambda::Function Properties: Handler: index.lambda_handler Runtime: python3.11 Role: !GetAtt CustomResourceLambdaRole.Arn Code: ZipFile: | import boto3 import cfnresponse import os s3 = boto3.client('s3') bucket_name = os.environ['BUCKET_NAME'] script_key = 'transform_json_to_parquet.py' def lambda_handler(event, context): try: if event['RequestType'] in ['Create', 'Update']: # 此处可替换为读取本地脚本、从其他S3桶拉取等逻辑 script_content = """ import sys from awsglue.transforms import * from awsglue.utils import getResolvedOptions from pyspark.context import SparkContext from awsglue.context import GlueContext from awsglue.job import Job args = getResolvedOptions(sys.argv, ['JOB_NAME', 's3_json_path', 's3_parquet_path']) sc = SparkContext() glueContext = GlueContext(sc) spark = glueContext.spark_session job = Job(glueContext) job.init(args['JOB_NAME'], args) dyf = glueContext.create_dynamic_frame.from_options( connection_type="s3", connection_options={"paths": [args['s3_json_path']]}, format="json" ) glueContext.write_dynamic_frame.from_options( frame=dyf, connection_type="s3", connection_options={"path": args['s3_parquet_path']}, format="parquet" ) job.commit() """ s3.put_object(Bucket=bucket_name, Key=script_key, Body=script_content.encode('utf-8')) cfnresponse.send(event, context, cfnresponse.SUCCESS, {}) except Exception as e: cfnresponse.send(event, context, cfnresponse.FAILED, {'Error': str(e)}) Environment: Variables: BUCKET_NAME: <S3-bucket-name-1> # 触发Lambda上传脚本的自定义资源 UploadGlueScriptResource: Type: AWS::CloudFormation::CustomResource Properties: ServiceToken: !GetAtt UploadGlueScriptLambda.Arn # 修改Glue Job,确保脚本上传完成后再创建 TransformJsonDataJob: Type: "AWS::Glue::Job" DependsOn: UploadGlueScriptResource Properties: Role: !Ref AWSGlueETLJobRole Name: "TransformJsonToParquet" Description: "Transform JSON to Parquet" Timeout: 5 WorkerType: G.1X NumberOfWorkers: 2 MaxRetries: 0 Command: "Name": "glueetl" "ScriptLocation" : !Sub s3://<S3-bucket-name-1>/transform_json_to_parquet.py DefaultArguments: "--s3_json_path" : !Sub s3://<S3-bucket-name-2>/ "--s3_parquet_path" : !Sub s3://<S3-bucket-name-3>/
方法3:用Shell脚本配合AWS CLI自动部署
如果不想修改CloudFormation模板,可以在部署栈前用脚本自动完成脚本上传和栈部署的流程:
# 先上传脚本到目标S3桶 aws s3 cp ./transform_json_to_parquet.py s3://<S3-bucket-name-1>/transform_json_to_parquet.py # 然后部署CloudFormation栈 aws cloudformation deploy --template-file your-template.yaml --stack-name your-stack-name --capabilities CAPABILITY_IAM
这种方法简单直接,适合脚本频繁变动的场景,只需将上述命令放到同一个Shell脚本中执行即可。
内容的提问来源于stack exchange,提问作者Pankesh Patel
相关产品推荐
相关产品推荐

