You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

JDK从11升级到17后Apache Ignite集群组建失败求助

问题描述

我正在基于Apache Ignite(从v2.8升级至v2.16.0)做一个简单POC项目:

  • JDK11(zulu11.37.17-ca-jdk11.0.6-win_x64)环境下,同一机器上的3个Ignite进程能正常组建集群:
    • 进程A:接收外部输入,调用进程B的服务并传递数据;
    • 进程B:发布服务供A调用,处理数据后传给进程C;
    • 进程C:接收B的数据处理,可调用A发布的服务。
  • 升级到JDK17(zulu17.48.15-ca-jdk17.0.10-win_x64)后,已按官方文档添加--add-opens等JVM参数,各进程能正常启动,但无法组建集群:进程A找不到B的服务,自身处于独立集群,即便B、C正常运行且配置一致。切回JDK11后,哪怕保留--add-opens参数,集群功能也恢复正常。

附上IgniteConfiguration.toString()内容:

IgniteConfiguration [igniteInstanceName=Ignite POC, pubPoolSize=8, svcPoolSize=null, callbackPoolSize=8, stripedPoolSize=8, sysPoolSize=8, mgmtPoolSize=4, dataStreamerPoolSize=8, utilityCachePoolSize=8, utilityCacheKeepAliveTime=60000, p2pPoolSize=2, qryPoolSize=8, buildIdxPoolSize=2, igniteHome=null, igniteWorkDir=null, mbeanSrv=null, nodeId=null, marsh=null, marshLocJobs=false, p2pEnabled=true, netTimeout=5000, netCompressionLevel=1, sndRetryDelay=1000, sndRetryCnt=3, metricsHistSize=10000, metricsUpdateFreq=2000, metricsExpTime=9223372036854775807, discoSpi=null, segPlc=USE_FAILURE_HANDLER, segResolveAttempts=2, waitForSegOnStart=true, allResolversPassReq=true, segChkFreq=10000, commSpi=null, evtSpi=null, colSpi=null, deploySpi=null, indexingSpi=null, addrRslvr=null, encryptionSpi=null, tracingSpi=null, clientMode=false, rebalanceThreadPoolSize=2, rebalanceTimeout=10000, rebalanceBatchesPrefetchCnt=3, rebalanceThrottle=0, rebalanceBatchSize=524288, txCfg=TransactionConfiguration [txSerEnabled=false, dfltIsolation=REPEATABLE_READ, dfltConcurrency=PESSIMISTIC, dfltTxTimeout=0, txTimeoutOnPartitionMapExchange=0, pessimisticTxLogSize=0, pessimisticTxLogLinger=10000, tmLookupClsName=null, txManagerFactory=null, useJtaSync=false], cacheSanityCheckEnabled=true, discoStartupDelay=60000, deployMode=SHARED, p2pMissedCacheSize=100, locHost=null, timeSrvPortBase=31100, timeSrvPortRange=100, failureDetectionTimeout=10000, sysWorkerBlockedTimeout=null, clientFailureDetectionTimeout=30000, metricsLogFreq=0, connectorCfg=ConnectorConfiguration [jettyPath=null, host=null, port=11211, noDelay=true, directBuf=false, sndBufSize=32768, rcvBufSize=32768, idleQryCurTimeout=600000, idleQryCurCheckFreq=60000, sndQueueLimit=0, selectorCnt=4, idleTimeout=7000, sslEnabled=false, sslClientAuth=false, sslFactory=null, portRange=100, threadPoolSize=8, msgInterceptor=null], odbcCfg=null, warmupClos=null, atomicCfg=AtomicConfiguration [seqReserveSize=1000, cacheMode=PARTITIONED, backups=1, aff=null, grpName=null], classLdr=null, sslCtxFactory=null, platformCfg=null, binaryCfg=null, memCfg=null, pstCfg=null, dsCfg=null, snapshotPath=snapshots, snapshotThreadPoolSize=4, activeOnStart=true, activeOnStartPropSetFlag=false, autoActivation=true, autoActivationPropSetFlag=false, clusterStateOnStart=null, sqlConnCfg=null, cliConnCfg=ClientConnectorConfiguration [host=null, port=10800, portRange=100, sockSndBufSize=0, sockRcvBufSize=0, tcpNoDelay=true, maxOpenCursorsPerConn=128, threadPoolSize=8, selectorCnt=4, idleTimeout=0, handshakeTimeout=10000, jdbcEnabled=true, odbcEnabled=true, thinCliEnabled=true, sslEnabled=false, useIgniteSslCtxFactory=true, sslClientAuth=false, sslCtxFactory=null, thinCliCfg=ThinClientConfiguration [maxActiveTxPerConn=100, maxActiveComputeTasksPerConn=0, sendServerExcStackTraceToClient=false], sesOutboundMsgQueueLimit=0], mvccVacuumThreadCnt=2, mvccVacuumFreq=5000, authEnabled=false, failureHnd=null, commFailureRslvr=null, sqlCfg=SqlConfiguration [longQryWarnTimeout=3000, dfltQryTimeout=0, sqlQryHistSize=1000, validationEnabled=false], asyncContinuationExecutor=null]
分析与解决建议

从配置和现象来看,JDK11正常、JDK17集群连不上,大概率是JDK17的模块限制或者网络默认行为变化导致的,给你列几个排查和解决的点:

1. 补全JDK17需要的所有JVM参数

你说只加了--add-opens,但Ignite 2.16.0在JDK17下需要的参数不止这些,得把所有必要的模块开放和权限设置加上:

--add-opens=java.base/java.lang=ALL-UNNAMED
--add-opens=java.base/java.nio=ALL-UNNAMED
--add-opens=java.base/java.util=ALL-UNNAMED
--add-opens=java.base/java.util.concurrent=ALL-UNNAMED
--add-opens=java.base/java.util.concurrent.atomic=ALL-UNNAMED
--add-opens=java.base/sun.nio.ch=ALL-UNNAMED
--add-opens=java.base/sun.security.provider=ALL-UNNAMED
--add-opens=java.management/javax.management=ALL-UNNAMED
--add-exports=java.base/sun.nio.ch=ALL-UNNAMED
--add-exports=java.base/sun.security.provider=ALL-UNNAMED
--add-exports=java.management/com.sun.jmx.mbeanserver=ALL-UNNAMED
--add-exports=jdk.internal.jvmstat/sun.jvmstat.monitor=ALL-UNNAMED
--add-exports=java.base/sun.reflect.generics.reflectiveObjects=ALL-UNNAMED

这些参数是Ignite在JDK17下正常运行的基础,少一个都可能导致内部反射失败,进而影响集群发现。

2. 手动配置集群发现机制

你的配置里discoSpi=null,也就是用默认的TcpDiscoverySpi,但JDK17对多播、本地网络探测的限制更严。建议显式配置TcpDiscoverySpi,指定本地IP和固定端口,别让它自动探测:

TcpDiscoverySpi discoSpi = new TcpDiscoverySpi();
TcpDiscoveryVmIpFinder ipFinder = new TcpDiscoveryVmIpFinder();
// 同一机器上三个进程用不同端口,比如47500、47501、47502
ipFinder.setAddresses(Arrays.asList("127.0.0.1:47500", "127.0.0.1:47501", "127.0.0.1:47502"));
discoSpi.setIpFinder(ipFinder);
igniteCfg.setDiscoverySpi(discoSpi);

手动指定IP和端口可以跳过自动多播发现,避免JDK17下多播权限或者网络接口探测的问题。

3. 给JDK17加网络权限

JDK17默认收紧了网络权限,得加个参数让Ignite能正常建立连接和通信:

--add-permissions=java.net.SocketPermission="*", "connect,resolve"

4. 关掉JDK17的预览特性

如果启动参数里有--enable-preview,赶紧去掉,别让这些不稳定特性影响Ignite运行。

5. 检查端口和防火墙

虽然是同一机器,但JDK17对端口绑定的逻辑可能变了,先确认三个进程的Ignite通信端口(默认47500-47509)没被占用,再看看Windows防火墙有没有拦截Ignite的请求。可以临时关闭防火墙测试,或者给Ignite进程添加防火墙例外。

6. 统一集群名称

你的配置没显式设置clusterName,默认值虽然一致,但建议手动加上统一的集群名,避免自动生成时出问题:

igniteCfg.setClusterName("Ignite-POC-Cluster");

确保三个节点都用这个配置,避免因为集群名不一致导致无法加入同一集群。

7. 查看日志找具体错误

启动时加个日志参数-DIGNITE_QUIET=false,把详细日志打出来,重点看集群发现阶段的错误,比如「Failed to join cluster」这类信息,直接就能定位问题根源。


内容的提问来源于stack exchange,提问作者J.M.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 07:45:57