EMR集群Presto/Trino JDBC密码认证配置与故障排查
我的使用场景十分简单:通过CDK部署了运行Presto的EMR集群,采用AWS Data Catalog作为元存储,集群仅使用默认用户执行查询。默认主用户为hadoop,我可通过该用户经JDBC连接集群执行查询,但当前无需输入密码即可成功建立连接。我查阅Presto官方文档后发现仅提及LDAP、Kerberos、基于文件的认证方式,我希望实现类似MySQL数据库的连接校验逻辑:连接时必须同时传入用户名和密码才可访问,但始终找不到设置hadoop用户密码的对应配置项,当前已完成的配置如下:
[ { "classification": "spark-hive-site", "configurationProperties": { "hive.metastore.client.factory.class": "com.amazonaws.glue.catalog.metastore.AWSGlueDataCatalogHiveClientFactory" } }, { "classification": "emrfs-site", "configurationProperties": { "fs.s3.maxConnections": "5000", "fs.s3.maxRetries": "200" } }, { "classification": "presto-connector-hive", "configurationProperties": { "hive.metastore.glue.datacatalog.enabled": "true", "hive.parquet.use-column-names": "true", "hive.max-partitions-per-writers": "7000000", "hive.table-statistics-enabled": "true", "hive.metastore.glue.max-connections": "20", "hive.metastore.glue.max-error-retries": "10", "hive.s3.use-instance-credentials": "true", "hive.s3.max-error-retries": "200", "hive.s3.max-client-retries": "100", "hive.s3.max-connections": "5000" } } ]
请问我可以通过哪个配置项设置hadoop用户的连接密码?Kerberos、LDAP、基于文件的认证方案对于这个简单使用场景来说过于复杂,我是否遗漏了显而易见的配置项?
在查阅大量官方文档并咨询AWS支持后,我决定切换使用Trino,但配置过程中遇到了更多问题,当前CDK部署的完整配置如下:
configurations: [ { "classification": "spark-hive-site", "configurationProperties": { "hive.metastore.client.factory.class": "com.amazonaws.glue.catalog.metastore.AWSGlueDataCatalogHiveClientFactory" } }, { "classification": "emrfs-site", "configurationProperties": { "fs.s3.maxConnections": "5000", "fs.s3.maxRetries": "200" } }, { "classification": "presto-connector-hive", "configurationProperties": { "hive.metastore.glue.datacatalog.enabled": "true", "hive.parquet.use-column-names": "true", "hive.max-partitions-per-writers": "7000000", "hive.table-statistics-enabled": "true", "hive.metastore.glue.max-connections": "20", "hive.metastore.glue.max-error-retries": "10", "hive.s3.use-instance-credentials": "true", "hive.s3.max-error-retries": "200", "hive.s3.max-client-retries": "100", "hive.s3.max-connections": "5000" } }, { "classification": "trino-config", "configurationProperties": { "query.max-memory-per-node": `${instanceMemory * 0.15}GB`, // 单节点查询最大内存,占节点内存25% "query.max-total-memory-per-node": `${instanceMemory * 0.5}GB`, // 单节点查询总内存上限,占节点内存50% "query.max-memory": `${instanceMemory * 0.5 * coreInstanceGroupNodeCount}GB`, // 集群查询总内存上限,占集群总内存50% "query.max-total-memory": `${instanceMemory * 0.8 * coreInstanceGroupNodeCount}GB`, // 集群查询总内存硬上限,占集群总内存80% "query.low-memory-killer.policy": "none", "task.concurrency": vcpuCount.toString(), "task.max-worker-threads": (vcpuCount * 4).toString(), "http-server.authentication.type": "PASSWORD", "http-server.http.enabled": "false", "internal-communication.shared-secret": "abcdefghijklnmopqrstuvwxyz", "http-server.https.enabled": "true", "http-server.https.port": "8443", "http-server.https.keystore.path": "/home/hadoop/fullCert.pem" } }, { "classification": "trino-password-authenticator", "configurationProperties": { "password-authenticator.name": "file", "file.password-file": "/home/hadoop/password.db", "file.refresh-period": "5s", "file.auth-token-cache.max-size": "1000" } } ]
我参考Trino官方TLS安全配置文档,采用直接加固Trino服务端的方案:即申请有效证书并添加到Trino协调器配置中。我已从公司获取内部通配符证书,包含证书文本、证书链、私钥三个部分。
参考PEM文件处理文档要求,我将三个部分合并为单个PEM文件,格式如下:
-----BEGIN RSA PRIVATE KEY----- Content of private key -----END RSA PRIVATE KEY----- -----BEGIN CERTIFICATE----- Content of certificate text -----END CERTIFICATE----- -----BEGIN CERTIFICATE----- First content of chain -----END CERTIFICATE----- -----BEGIN CERTIFICATE----- Second content of chain -----END CERTIFICATE-----
我通过引导操作将合并后的证书文件部署到所有节点,满足Trino协调器TLS配置要求,对应配置项为:
'http-server.https.enabled': 'true', 'http-server.https.port': '8443', 'http-server.https.keystore.path': '/home/hadoop/fullCert.pem',
我已确认证书文件成功部署到所有节点,随后参考密码文件认证文档完成了密码认证配置,该部分功能已验证生效:在主节点上使用错误密码通过trino-cli连接时会返回凭证错误。
当前遇到的问题
使用带
--insecure参数的trino-cli连接本地8446端口执行查询时,报错提示活动工作节点不足,等待5分钟仍未检测到至少1个可用工作节点,实际可用工作节点数为0:[hadoop@ip-10-0-10-245 ~]$ trino-cli --server https://localhost:8446 --catalog awsdatacatalog --user hadoop --password --insecure trino> select 1; Query 20220701_201620_00001_9nksi failed: Insufficient active worker nodes. Waited 5.00m for at least 1 workers, but only 0 workers are active查看
/var/log/trino/server.log日志,存在节点状态获取失败、服务持续通告报错的问题:2022-07-01T21:30:12.966Z WARN http-client-node-manager-51 io.trino.metadata.RemoteNodeState Error fetching node state from https://ip-10-0-10-245.ec2.internal:8446/v1/info/state: Failed communicating with server: https://ip-10-0-10-245.ec2.internal:8446/v1/info/state 2022-07-01T21:30:13.902Z ERROR Announcer-0 io.airlift.discovery.client.Announcer Service announcement failed after 8.11ms. Next request will happen within 1000.00ms 2022-07-01T21:30:14.913Z ERROR Announcer-1 io.airlift.discovery.client.Announcer Service announcement failed after 10.35ms. Next request will happen within 1000.00ms 2022-07-01T21:30:15.921Z ERROR Announcer-3 io.airlift.discovery.client.Announcer Service announcement failed after 8.40ms. Next request will happen within 1000.00ms 2022-07-01T21:30:16.930Z ERROR Announcer-0 io.airlift.discovery.client.Announcer Service announcement failed after 8.59ms. Next request will happen within 1000.00ms 2022-07-01T21:30:17.938Z ERROR Announcer-1 io.airlift.discovery.client.Announcer Service announcement failed after 8.36ms. Next request will happen within 1000.00ms不带
--insecure参数通过trino-cli连接时,报SSL握手失败错误,提示无法找到请求目标的有效证书路径:[hadoop@ip-10-0-10-245 ~]$ trino-cli --server https://localhost:8446 --catalog awsdatacatalog --user hadoop --password trino> select 1; Error running command: javax.net.ssl.SSLHandshakeException: PKIX path building failed: sun.security.provider.certpath.SunCertPathBuilderException: unable to find valid certification path to requested target trino>
我已按照AWS EMR加密配置文档要求将PEM证书文件作为资产上传到S3,希望得到上述问题的解决方案,为何简单的密码认证配置流程会如此复杂?
针对最初Presto密码配置问题的说明
EMR自带的旧版Presto(Trino的前身分支)不存在单配置项设置默认用户密码的能力,你没有遗漏配置项:Presto原生设计就没有内置类似MySQL的用户名密码校验逻辑,所有密码类认证必须依赖LDAP、Kerberos或者外置文件认证插件,没有开箱即用的单用户密码设置项。如果不想部署LDAP/Kerberos这类重认证方案,切换Trino使用文件密码认证是最小成本的实现路径,也就是你当前选择的方案。
Trino问题1:工作节点数为0、服务通告失败
这里有两个配置错误直接导致该问题:
- 端口配置不匹配:你在
trino-config中配置的HTTPS端口为8443,但实际连接和日志中访问的端口是8446,端口不一致会导致节点之间无法正常通信。要么统一将所有节点的http-server.https.port配置为你实际使用的8446,要么连接时使用配置中指定的8443端口,保证协调器监听端口、工作节点通告端口、客户端连接端口三者完全一致。 - 内部通信TLS配置缺失:你仅开启了对外服务的HTTPS,同时禁用了HTTP端口,但没有配置集群内部节点间通信的TLS参数,工作节点和协调器之间默认仍尝试走HTTP通信,自然无法建立连接。需要补充以下配置项:
另外需要确认所有节点(含工作节点)上的证书文件对Trino进程可读:你当前将证书放在# 开启内部通信HTTPS 'internal-communication.https.enabled': 'true', 'internal-communication.https.keystore.path': '/home/hadoop/fullCert.pem', # 配置内部通信信任证书,适配内部签发的证书 'internal-communication.https.truststore.path': '/home/hadoop/fullCert.pem', 'internal-communication.https.truststore.password': '', # 若证书设置了密码则填写对应值,无密码留空即可/home/hadoop目录下,默认权限可能仅允许hadoop用户访问,执行以下命令修正权限即可:chmod 644 /home/hadoop/fullCert.pem && chown trino:trino /home/hadoop/fullCert.pem
Trino问题2:无--insecure参数时SSL握手失败
该问题的根因是Java客户端默认信任库不包含你使用的内部证书的签发CA:trino-cli基于Java开发,默认仅信任JDK内置的公共CA根证书,无法识别你们公司内部签发的证书,因此校验失败。有两种解决方式:
- 测试环境可直接加
--insecure参数跳过客户端侧证书校验,不影响密码认证功能的正常使用。 - 生产环境建议将公司内部证书的根CA、中间CA证书导入所有节点的JRE默认信任库,默认信任库路径为
/etc/alternatives/jre/lib/security/cacerts,默认密码为changeit,导入命令参考:
导入完成后重启所有节点的Trino服务,后续连接无需加keytool -importcert -alias internal-ca-root -file /path/to/your/root-ca.crt -keystore /etc/alternatives/jre/lib/security/cacerts -storepass changeit -noprompt--insecure参数即可正常完成TLS握手。
注意:你当前配置中硬编码的
internal-communication.shared-secret是弱密钥,生产环境请替换为随机生成的32位以上长度的随机字符串,避免内部通信被伪造篡改。
内容的提问来源于stack exchange,提问作者rodrigocf

