Hudi DeltaStreamer同步AWS Glue数据目录仅同步数据库未同步表
Hudi DeltaStreamer同步AWS Glue数据目录异常:仅同步数据库,表未同步
问题概述
该问题与Stack Overflow上「运行Hudi DeltaStreamer成功但无法同步AWS Glue数据目录」的场景类似,核心差异是仅能同步数据库,无法同步表。
提交的Spark命令
spark-submit \ --conf spark.hadoop.hive.metastore.client.factory.class=com.amazonaws.glue.catalog.metastore.AWSGlueDataCatalogHiveClientFactory \ --deploy-mode cluster \ --jars /usr/lib/spark/external/lib/spark-avro.jar,/usr/lib/hudi/hudi-spark-bundle.jar,/usr/lib/hudi/hudi-utilities-bundle.jar,/usr/lib/hudi/cli/lib/aws-java-sdk-glue-1.12.397.jar,/usr/lib/hive/auxlib/aws-glue-datacatalog-hive3-client.jar,/usr/lib/hadoop/hadoop-aws.jar,/usr/lib/hadoop/hadoop-aws-3.3.3-amzn-2.jar \ --conf spark.sql.catalogImplementation=hive \ --conf spark.sql.catalog.spark_catalog=org.apache.spark.sql.hudi.catalog.HoodieCatalog \ --conf spark.sql.extensions=org.apache.spark.sql.hudi.HoodieSparkSessionExtension \ --class org.apache.hudi.utilities.deltastreamer.HoodieDeltaStreamer /usr/lib/hudi/hudi-utilities-slim-bundle.jar \ --table-type COPY_ON_WRITE \ --source-class org.apache.hudi.utilities.sources.AvroDFSSource \ --source-ordering-field id \ --target-base-path s3a://my-bucket/data/my_database/my_target_table/ \ --sync-tool-classes org.apache.hudi.aws.sync.AwsGlueCatalogSyncTool \ --props file:///etc/hudi/conf/hudi-defaults.conf \ --target-table my_target_table \ --schemaprovider-class org.apache.hudi.utilities.schema.SchemaRegistryProvider \ --enable-sync \ --enable-hive-sync
执行后的现象
通过Beeline连接Hive后,数据库已同步,但表未同步:
beeline -u jdbc:hive2://ip-1-1-1-1:10000 Connecting to jdbc:hive2://ip-1-1-1:10000 show databases; +-----------------------------------+ | database_name | +-----------------------------------+ | my_database | +-----------------------------------+ show tables; +----------------------------------------------------+ | tab_name | +----------------------------------------------------+ | | +----------------------------------------------------+
排查与解决要点
- 检查
hudi-defaults.conf配置:确保hoodie.datasource.hive_sync.database设为my_database,hoodie.datasource.hive_sync.table设为my_target_table,hoodie.datasource.hive_sync.enable设为true。 - 验证IAM权限:运行DeltaStreamer的角色需具备Glue的
CreateTable、UpdateTable权限,同时对S3目标路径有读写权限。 - 修正命令参数:
--target-table需指定完整库表名(如my_database.my_target_table),避免仅传表名导致同步路径不匹配。 - 查看集群日志:搜索Spark日志中Glue同步相关条目,定位是否存在Schema不匹配、权限报错等具体问题。
- 确认版本兼容性:确保Hudi版本与AWS Glue Data Catalog客户端版本兼容,避免因版本冲突导致同步失败。
内容的提问来源于stack exchange,提问作者Will
相关产品推荐
相关产品推荐

