如何在AWS Glue 4中修改Hudi表的存储位置?
问题描述
我使用AWS Glue 4,现有一张存储在s3://aws-amazon-com/Customer/的Hudi Customer表,希望将其存储位置修改为s3://aws-amazon-com/CustomerUpdated/。使用的依赖jar包包括:
- hudi-spark3-bundle_2.12-0.12.1.jar
- calcite-core-1.16.0.jar
- libfb303-0.9.3.jar
我编写了以下Scala代码尝试将数据写入新位置:
val partitionColumnName: String = "year" val hudiTableName: String = "Customer" val preCombineKey: String = "id" val recordKey = "id" val tablePath = "s3://aws-amazon-com/Customer/" val databaseName="consumer_bureau" val hudiCommonOptions: Map[String, String] = Map( "hoodie.table.name" -> hudiTableName, "hoodie.datasource.write.keygenerator.class" -> "org.apache.hudi.keygen.ComplexKeyGenerator", "hoodie.datasource.write.precombine.field" -> preCombineKey, "hoodie.datasource.write.recordkey.field" -> recordKey, "hoodie.datasource.write.operation" -> "bulk_insert", //"hoodie.datasource.write.operation" -> "upsert", "hoodie.datasource.write.row.writer.enable" -> "true", "hoodie.datasource.write.reconcile.schema" -> "true", "hoodie.datasource.write.partitionpath.field" -> partitionColumnName, "hoodie.datasource.write.hive_style_partitioning" -> "true", // "hoodie.bulkinsert.shuffle.parallelism" -> "2000", // "hoodie.upsert.shuffle.parallelism" -> "400", "hoodie.datasource.hive_sync.enable" -> "true", "hoodie.datasource.hive_sync.table" -> hudiTableName, "hoodie.datasource.hive_sync.database" -> databaseName, "hoodie.datasource.hive_sync.partition_fields" -> partitionColumnName, "hoodie.datasource.hive_sync.partition_extractor_class" -> "org.apache.hudi.hive.MultiPartKeysValueExtractor", "hoodie.datasource.hive_sync.use_jdbc" -> "false", "hoodie.combine.before.upsert" -> "true", "hoodie.index.type" -> "BLOOM", "spark.hadoop.parquet.avro.write-old-list-structure" -> "false", DataSourceWriteOptions.TABLE_TYPE.key() -> "COPY_ON_WRITE" ) val df=Seq((1,"Mark",1990),(2,"Martin",2009)).toDF("id","name","year") df.write.format("org.apache.hudi") .options(hudiCommonOptions) .mode(SaveMode.Append) .save(tablelocation) val tablelocationUpdated="s3://eec-aws-uk-ukidcibatchanalytics-prod-hudi-replication/consumer_bureau/production/CustomerUpdated/" df.write.format("org.apache.hudi") //writng to new location .options(hudiCommonOptions) .mode(SaveMode.Append) .save(tablelocationUpdated)
但通过Athena查询时,Customer表仍指向旧存储路径s3://aws-amazon-com/Customer/,而非预期的新路径。请问能否通过AWS Glue或AWS Lambda实现Hudi表的存储位置变更?
解决方案
你的问题核心在于:直接写入新路径只是创建了一个新的Hudi表,但Glue Data Catalog中原有Customer表的元数据没更新,所以Athena还是指向旧路径。下面是两种可行的实现方式:
方法一:通过AWS Glue修改表存储位置
步骤1:迁移Hudi表数据到新路径
必须完整迁移旧路径的所有Hudi数据(包括.hoodie元数据文件夹),否则新路径的表会丢失事务能力,有两种方式可选:
- Glue作业迁移:编写Scala/PySpark作业读取旧表全量数据,写入新路径,示例Scala代码:
val oldTablePath = "s3://aws-amazon-com/Customer/" val newTablePath = "s3://aws-amazon-com/CustomerUpdated/" val hudiReadOptions = Map("hoodie.datasource.read.path" -> oldTablePath) val hudiDF = spark.read.format("org.apache.hudi").options(hudiReadOptions).load() // 复用原有配置,替换存储路径 val updatedOptions = hudiCommonOptions + ("hoodie.datasource.write.path" -> newTablePath) hudiDF.write.format("org.apache.hudi").options(updatedOptions).mode(SaveMode.Overwrite).save(newTablePath)
- S3命令行同步:用AWS CLI直接同步所有文件:
aws s3 sync s3://aws-amazon-com/Customer/ s3://aws-amazon-com/CustomerUpdated/
步骤2:更新Glue表元数据
数据迁移完成后,更新Glue Data Catalog中的表存储位置:
- 打开AWS Glue控制台,进入
consumer_bureau数据库,找到Customer表 - 点击「编辑表」,将「存储位置」修改为
s3://aws-amazon-com/CustomerUpdated/ - 保存后,Athena查询就会指向新路径
如果需要自动化操作,可以用Glue Python作业调用boto3修改:
import boto3 glue_client = boto3.client('glue') response = glue_client.update_table( DatabaseName='consumer_bureau', TableInput={ 'Name': 'Customer', 'StorageDescriptor': { 'Location': 's3://aws-amazon-com/CustomerUpdated/', # 保留原有StorageDescriptor的其他配置,比如列信息、Serde配置等 } } )
方法二:通过AWS Lambda实现自动化变更
Lambda可以一键完成「数据迁移+元数据更新」,适合定时或触发式执行的场景:
- 创建Lambda函数,选择Python 3.x运行时,配置IAM权限(需要S3读写、Glue表修改权限)
- 编写Lambda代码实现核心逻辑:
import boto3 import subprocess def lambda_handler(event, context): # 配置参数 old_s3_path = "s3://aws-amazon-com/Customer/" new_s3_path = "s3://aws-amazon-com/CustomerUpdated/" database_name = "consumer_bureau" table_name = "Customer" # 同步S3数据 subprocess.run(["aws", "s3", "sync", old_s3_path, new_s3_path], check=True) # 更新Glue表元数据(先获取原有配置,避免覆盖其他字段) glue_client = boto3.client('glue') table = glue_client.get_table(DatabaseName=database_name, Name=table_name)['Table'] table['StorageDescriptor']['Location'] = new_s3_path # 移除系统自动生成的字段,避免更新报错 for key in ['CreatedBy', 'CreateTime', 'UpdateTime']: table.pop(key, None) glue_client.update_table( DatabaseName=database_name, TableInput=table ) return {"statusCode": 200, "message": "Hudi表存储位置更新成功"}
注意事项
- 迁移时必须同步
.hoodie文件夹,否则新表无法使用Hudi的增量、更新等核心功能 - 若为COPY_ON_WRITE类型的Hudi表,迁移后建议验证新旧路径的记录数,确保数据完整
- Glue作业迁移时,要确保作业依赖的jar包包含指定的Hudi依赖,且Glue 4.0与Hudi 0.12.1版本兼容
内容的提问来源于stack exchange,提问作者gaurav mathur
相关产品推荐
相关产品推荐

