Spark自定义JDBC源连接Kerberos认证Hive报错的解决咨询
针对你遇到的Driver端Kerberos认证成功,但Executor执行JDBC连接时抛出sqlexception:could not open client for any of server uri is zookeeper:null的问题,按以下步骤排查解决:
1. 确保Executor加载完整的Hadoop/Hive配置
Driver本地的配置不会自动同步到Executor,必须手动传递关键配置文件:
- 提交Spark任务时用
--files参数指定需要的配置文件,让Executor能获取到正确的Zookeeper和Kerberos配置:spark-submit --files /etc/hive/conf/hive-site.xml,/etc/hadoop/conf/core-site.xml,/etc/hadoop/conf/hdfs-site.xml,/etc/krb5.conf \ --class your.main.class \ your-spark-job.jar - 或者配置
spark.driver.extraClassPath和spark.executor.extraClassPath,指向存放这些配置文件的目录,确保Executor能读取到。
2. 让Executor获取Kerberos凭证
yarn-client模式下,Driver在本地,Executor在集群节点,需要把Kerberos凭证分发到Executor:
- 开启Spark的Hive凭证自动传递:
spark-submit --conf spark.yarn.security.credentials.hive.enabled=true \ --conf spark.security.credentials.hive.enabled=true \ ... - 确保Driver端的Kerberos缓存文件(默认是
/tmp/krb5cc_<uid>)能被YARN正确传递,或者在任务提交前通过kinit生成有效凭证,避免Executor无凭证可用。
3. 修正JDBC连接URL
自定义JDBC数据源的连接URL必须包含正确的Zookeeper集群地址,格式如下:
jdbc:hive2://zk1:2181,zk2:2181,zk3:2181/;serviceDiscoveryMode=zooKeeper;zooKeeperNamespace=hiveserver2
确认<zookeeper_quorum>是集群中Zookeeper节点的真实地址,不要留空或配置错误。
4. 正确配置Executor的JVM参数
之前用的spark.javaextraoption参数错误,应该用spark.executor.extraJavaOptions传递Kerberos相关JVM参数:
spark-submit --conf spark.executor.extraJavaOptions="-Djava.security.krb5.conf=/etc/krb5.conf -Dsun.security.krb5.debug=true" \ ...
开启sun.security.krb5.debug=true可以输出详细的Kerberos认证日志,方便定位问题。
5. 检查HiveServer2的Zookeeper注册状态
登录Zookeeper节点,用命令验证HiveServer2是否正确注册:
zkCli.sh ls /hiveserver2
如果没有返回HiveServer2的节点信息,说明HiveServer2未正确注册,检查hive-site.xml中的hive.zookeeper.quorum、hive.zookeeper.client.port配置,然后重启HiveServer2。
6. 确认Executor节点的证书信任
虽然所有节点已部署证书,仍需确保Executor节点的JVM信任存储包含Hive和Zookeeper的证书,避免SSL连接失败:
keytool -import -trustcacerts -alias hive_cert -file /path/to/hive.crt -keystore $JAVA_HOME/jre/lib/security/cacerts
输入默认密码changeit完成证书导入,重启Executor节点的YARN nodemanager服务。
内容的提问来源于stack exchange,提问作者suhenhappy

