EMR 5.28.0开启Glue catalog时Hudi同步Hive出现jdbc连接错误如何解决
问题排查与解决方案
错误根因
该报错是因为你已将AWS Glue设置为元数据目录,但Hudi的hive-sync功能默认仍尝试连接本地HiveServer2(默认地址jdbc:hive2://localhost:10000/),而EMR集群如果未启动HiveServer2服务、或未配置Hudi对接Glue元数据,就会触发连接失败异常。
解决步骤
- 确认集群IAM权限
确保EMR集群以及Jupyter Notebook的执行角色具备Glue表读写权限、S3目标路径的读写权限,可通过IAM控制台检查对应角色是否关联AWSGlueServiceRole类权限策略。 - 新增Hudi对接Glue元数据的配置
直接在你的hudiOptions中新增Glue同步参数,开启后Hudi会直接调用Glue API同步元数据,不需要连接本地HiveServer2,是该场景下的最优解:# 新增到hudiOptions中的配置 'hoodie.datasource.hive_sync.use_glue': 'true', 'hoodie.datasource.hive_sync.ignore_glue_skip_archive': 'true' - (可选)若坚持使用原生Hive同步模式
登录EMR控制台确认HiveServer2服务已正常启动,也可登录主节点执行命令sudo systemctl status hive-server2查看服务状态,未启动则手动启动该服务即可。
修正后完整配置示例
# Specify common DataSourceWriteOptions in the single hudiOptions variable hudiOptions = { 'hoodie.table.name': 'my_hudi_table', 'hoodie.datasource.write.recordkey.field': 'id', 'hoodie.datasource.write.partitionpath.field': 'creation_date', 'hoodie.datasource.write.precombine.field': 'last_update_time', 'hoodie.datasource.hive_sync.enable': 'true', 'hoodie.datasource.hive_sync.table': 'my_hudi_table', 'hoodie.datasource.hive_sync.partition_fields': 'creation_date', 'hoodie.datasource.hive_sync.partition_extractor_class': 'org.apache.hudi.hive.MultiPartKeysValueExtractor', # 新增Glue同步配置 'hoodie.datasource.hive_sync.use_glue': 'true', 'hoodie.datasource.hive_sync.ignore_glue_skip_archive': 'true' }
内容的提问来源于stack exchange,提问作者dytyniak
相关产品推荐
相关产品推荐

