You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在AWS Glue 4中修改Hudi表的存储位置?

问题描述

我使用AWS Glue 4,现有一张存储在s3://aws-amazon-com/Customer/的Hudi Customer表,希望将其存储位置修改为s3://aws-amazon-com/CustomerUpdated/。使用的依赖jar包包括:

  • hudi-spark3-bundle_2.12-0.12.1.jar
  • calcite-core-1.16.0.jar
  • libfb303-0.9.3.jar

我编写了以下Scala代码尝试将数据写入新位置:

val partitionColumnName: String = "year"
val hudiTableName: String = "Customer"
val preCombineKey: String = "id"
val recordKey = "id"
val tablePath = "s3://aws-amazon-com/Customer/"
val databaseName="consumer_bureau"

val hudiCommonOptions: Map[String, String] = Map(
    "hoodie.table.name" -> hudiTableName,
    "hoodie.datasource.write.keygenerator.class" -> "org.apache.hudi.keygen.ComplexKeyGenerator",
    "hoodie.datasource.write.precombine.field" -> preCombineKey,
    "hoodie.datasource.write.recordkey.field" -> recordKey,
    "hoodie.datasource.write.operation" -> "bulk_insert",
    //"hoodie.datasource.write.operation" -> "upsert",
    "hoodie.datasource.write.row.writer.enable" -> "true",
    "hoodie.datasource.write.reconcile.schema" -> "true",
    "hoodie.datasource.write.partitionpath.field" -> partitionColumnName,
    "hoodie.datasource.write.hive_style_partitioning" -> "true",
    // "hoodie.bulkinsert.shuffle.parallelism" -> "2000",
    //  "hoodie.upsert.shuffle.parallelism" -> "400",
    "hoodie.datasource.hive_sync.enable" -> "true",
    "hoodie.datasource.hive_sync.table" -> hudiTableName,
    "hoodie.datasource.hive_sync.database" -> databaseName,
    "hoodie.datasource.hive_sync.partition_fields" -> partitionColumnName,
    "hoodie.datasource.hive_sync.partition_extractor_class" -> "org.apache.hudi.hive.MultiPartKeysValueExtractor",
    "hoodie.datasource.hive_sync.use_jdbc" -> "false",
    "hoodie.combine.before.upsert" -> "true",
    "hoodie.index.type" -> "BLOOM",
    "spark.hadoop.parquet.avro.write-old-list-structure" -> "false",
    DataSourceWriteOptions.TABLE_TYPE.key() -> "COPY_ON_WRITE"
  )
  
val df=Seq((1,"Mark",1990),(2,"Martin",2009)).toDF("id","name","year")
  
df.write.format("org.apache.hudi")
.options(hudiCommonOptions)
.mode(SaveMode.Append)
.save(tablelocation)
    
val tablelocationUpdated="s3://eec-aws-uk-ukidcibatchanalytics-prod-hudi-replication/consumer_bureau/production/CustomerUpdated/"
   
df.write.format("org.apache.hudi") //writng to new location
.options(hudiCommonOptions)
.mode(SaveMode.Append)
.save(tablelocationUpdated)

但通过Athena查询时,Customer表仍指向旧存储路径s3://aws-amazon-com/Customer/,而非预期的新路径。请问能否通过AWS Glue或AWS Lambda实现Hudi表的存储位置变更?


解决方案

你的问题核心在于:直接写入新路径只是创建了一个新的Hudi表,但Glue Data Catalog中原有Customer表的元数据没更新,所以Athena还是指向旧路径。下面是两种可行的实现方式:

方法一:通过AWS Glue修改表存储位置

步骤1:迁移Hudi表数据到新路径

必须完整迁移旧路径的所有Hudi数据(包括.hoodie元数据文件夹),否则新路径的表会丢失事务能力,有两种方式可选:

  • Glue作业迁移:编写Scala/PySpark作业读取旧表全量数据,写入新路径,示例Scala代码:
val oldTablePath = "s3://aws-amazon-com/Customer/"
val newTablePath = "s3://aws-amazon-com/CustomerUpdated/"
val hudiReadOptions = Map("hoodie.datasource.read.path" -> oldTablePath)
val hudiDF = spark.read.format("org.apache.hudi").options(hudiReadOptions).load()

// 复用原有配置,替换存储路径
val updatedOptions = hudiCommonOptions + ("hoodie.datasource.write.path" -> newTablePath)
hudiDF.write.format("org.apache.hudi").options(updatedOptions).mode(SaveMode.Overwrite).save(newTablePath)
  • S3命令行同步:用AWS CLI直接同步所有文件:
aws s3 sync s3://aws-amazon-com/Customer/ s3://aws-amazon-com/CustomerUpdated/

步骤2:更新Glue表元数据

数据迁移完成后,更新Glue Data Catalog中的表存储位置:

  1. 打开AWS Glue控制台,进入consumer_bureau数据库,找到Customer表
  2. 点击「编辑表」,将「存储位置」修改为s3://aws-amazon-com/CustomerUpdated/
  3. 保存后,Athena查询就会指向新路径

如果需要自动化操作,可以用Glue Python作业调用boto3修改:

import boto3

glue_client = boto3.client('glue')

response = glue_client.update_table(
    DatabaseName='consumer_bureau',
    TableInput={
        'Name': 'Customer',
        'StorageDescriptor': {
            'Location': 's3://aws-amazon-com/CustomerUpdated/',
            # 保留原有StorageDescriptor的其他配置,比如列信息、Serde配置等
        }
    }
)

方法二:通过AWS Lambda实现自动化变更

Lambda可以一键完成「数据迁移+元数据更新」,适合定时或触发式执行的场景:

  1. 创建Lambda函数,选择Python 3.x运行时,配置IAM权限(需要S3读写、Glue表修改权限)
  2. 编写Lambda代码实现核心逻辑:
import boto3
import subprocess

def lambda_handler(event, context):
    # 配置参数
    old_s3_path = "s3://aws-amazon-com/Customer/"
    new_s3_path = "s3://aws-amazon-com/CustomerUpdated/"
    database_name = "consumer_bureau"
    table_name = "Customer"
    
    # 同步S3数据
    subprocess.run(["aws", "s3", "sync", old_s3_path, new_s3_path], check=True)
    
    # 更新Glue表元数据(先获取原有配置,避免覆盖其他字段)
    glue_client = boto3.client('glue')
    table = glue_client.get_table(DatabaseName=database_name, Name=table_name)['Table']
    table['StorageDescriptor']['Location'] = new_s3_path
    
    # 移除系统自动生成的字段,避免更新报错
    for key in ['CreatedBy', 'CreateTime', 'UpdateTime']:
        table.pop(key, None)
    
    glue_client.update_table(
        DatabaseName=database_name,
        TableInput=table
    )
    
    return {"statusCode": 200, "message": "Hudi表存储位置更新成功"}

注意事项

  • 迁移时必须同步.hoodie文件夹,否则新表无法使用Hudi的增量、更新等核心功能
  • 若为COPY_ON_WRITE类型的Hudi表,迁移后建议验证新旧路径的记录数,确保数据完整
  • Glue作业迁移时,要确保作业依赖的jar包包含指定的Hudi依赖,且Glue 4.0与Hudi 0.12.1版本兼容

内容的提问来源于stack exchange,提问作者gaurav mathur

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 07:57:02