You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark YARN模式启动失败:Java网关及ApplicationMaster超时问题

PySpark on YARN 错误排查与解决

环境背景

Ubuntu 18系统,1主2从Hadoop集群,Hadoop与YARN配置完成,JAVA_HOME、YARN_CONF_DIR、HADOOP_CONF_DIR已在~/.bashrc中配置,从节点yarn-site.xml指定yarn.resourcemanager.hostname为hadoop-master,YARN仪表盘可通过http://hadoop-master:8088/cluster正常访问。本地模式运行PySpark无异常,切换YARN模式后出现以下两类错误:


错误1:设置master为yarn_client时触发Java gateway退出

错误栈:

RuntimeError                              Traceback (most recent call last)
<ipython-input-3-2a71daf20656> in <module>
      1 conf = SparkConf().setAppName("Spark_hadoop").setMaster("yarn_client").set("spark.executor.memory","5g")
----> 2 sc = SparkContext(conf=conf)
      3 sc=SparkContext.getOrCreate(conf=create_spark_conf().setMaster("local[4]").set("spark.driver.memory","8g").set("spark.executor.memory", '8g').set('spark.executor.cores', 4))
      4 sc.setLogLevel("ERROR")
      5 sqlContext = SQLContext(sc)

~/.local/lib/python3.6/site-packages/pyspark/context.py in __init__(self, master, appName, sparkHome, pyFiles, environment, batchSize, serializer, conf, gateway, jsc, profiler_cls)
    142                 " is not allowed as it is a security risk.")
    143 
--> 144         SparkContext._ensure_initialized(self, gateway=gateway, conf=conf)
    145         try:
    146             self._do_init(master, appName, sparkHome, pyFiles, environment, batchSize, serializer,

~/.local/lib/python3.6/site-packages/pyspark/context.py in _ensure_initialized(cls, instance, gateway, conf)
    337         with SparkContext._lock:
    338             if not SparkContext._gateway:
--> 339                 SparkContext._gateway = gateway or launch_gateway(conf)
    340                 SparkContext._jvm = SparkContext._gateway.jvm
    341 

~/.local/lib/python3.6/site-packages/pyspark/java_gateway.py in launch_gateway(conf, popen_kwargs)
    106 
    107             if not os.path.isfile(conn_info_file):
--> 108                 raise RuntimeError("Java gateway process exited before sending its port number")
    109 
    110             with open(conn_info_file, "rb") as info:

RuntimeError: Java gateway process exited before sending its port number

解决步骤:

  • 修正master参数格式:Spark官方已统一使用yarn作为YARN模式的master参数,yarn_client或旧版yarn-client格式已被弃用。如需指定client模式,需通过spark.submit.deployMode参数配置:
    conf = SparkConf().setAppName("Spark_hadoop") \
                      .setMaster("yarn") \
                      .set("spark.submit.deployMode", "client") \
                      .set("spark.executor.memory", "5g")
    
  • 验证Java环境一致性:在Jupyter Notebook中执行!echo $JAVA_HOME和!java -version,确保与Hadoop/YARN使用的Java版本、路径完全一致。
  • 清理临时文件:删除/tmp目录下前缀为spark-的临时文件,重启Jupyter后重试。

错误2:改为yarn模式后触发ApplicationMaster超时

错误栈:

Py4JJavaError                             Traceback (most recent call last)
<ipython-input-7-2ee19c87679b> in <module>
      2 findspark.init()
      3 conf = SparkConf().setAppName("Spark_hadoop").setMaster("yarn").set("spark.executor.memory","5g")
----> 4 sc = SparkContext(conf=conf)
      5 sc.setLogLevel("ERROR")
      6 sqlContext = SQLContext(sc)

~/.local/lib/python3.6/site-packages/pyspark/context.py in __init__(self, master, appName, sparkHome, pyFiles, environment, batchSize, serializer, conf, gateway, jsc, profiler_cls)
    145         try:
    146             self._do_init(master, appName, sparkHome, pyFiles, environment, batchSize, serializer,
--> 147                           conf, jsc, profiler_cls)
    148         except:
    149             # If an error occurs, clean up in order to allow future SparkContext creation:

~/.local/lib/python3.6/site-packages/pyspark/context.py in _do_init(self, master, appName, sparkHome, pyFiles, environment, batchSize, serializer, conf, jsc, profiler_cls)
    207 
    208         # Create the Java SparkContext through Py4J
--> 209         self._jsc = jsc or self._initialize_context(self._conf._jconf)
    210         # Reset the SparkConf to the one actually used by the SparkContext in JVM.
    211         self._conf = SparkConf(_jconf=self._jsc.sc().conf())

~/.local/lib/python3.6/site-packages/pyspark/context.py in _initialize_context(self, jconf)
    327         Initialize SparkContext in function to allow subclass specific initialization
    328         """
--> 329         return self._jvm.JavaSparkContext(jconf)
    330 
    331     @classmethod

~/.local/lib/python3.6/site-packages/py4j/java_gateway.py in __call__(self, *args)
   1584         answer = self._gateway_client.send_command(command)
   1585         return_value = get_return_value(
-> 1586             answer, self._gateway_client, None, self._fqn)
   1587 
   1588         for temp_arg in temp_args:

~/.local/lib/python3.6/site-packages/pyspark/sql/utils.py in deco(*a, **kw)
    109     def deco(*a, **kw):
    110         try:
--> 111             return f(*a, **kw)
    112         except py4j.protocol.Py4JJavaError as e:
    113             converted = convert_exception(e.java_exception)

~/.local/lib/python3.6/site-packages/py4j/protocol.py in get_return_value(answer, gateway_client, target_id, name)
    326                 raise Py4JJavaError(
    327                     "An error occurred while calling {0}{1}{2}.\n".
-> 328                     format(target_id, ".", name), value)
    329             else:
    330                 raise Py4JError(

Py4JJavaError: An error occurred while calling None.org.apache.spark.api.java.JavaSparkContext.
: org.apache.spark.SparkException: Application application_1675283388270_0012 failed 2 times due to ApplicationMaster for attempt appattempt_1675283388270_0012_000002 timed out. Failing the application.
    at org.apache.spark.scheduler.cluster.YarnClientSchedulerBackend.waitForApplication(YarnClientSchedulerBackend.scala:98)
    at org.apache.spark.scheduler.cluster.YarnClientSchedulerBackend.start(YarnClientSchedulerBackend.scala:65)
    at org.apache.spark.scheduler.TaskSchedulerImpl.start(TaskSchedulerImpl.scala:222)
    at org.apache.spark.SparkContext.<init>(SparkContext.scala:585)
    at org.apache.spark.api.java.JavaSparkContext.<init>(JavaSparkContext.scala:58)
    at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)
    at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)
    at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)
    at java.lang.reflect.Constructor.newInstance(Constructor.java:423)
    at py4j.reflection.MethodInvoker.invoke(MethodInvoker.java:247)
    at py4j.reflection.ReflectionEngine.invoke(ReflectionEngine.java:357)
    at py4j.Gateway.invoke(Gateway.java:238)
    at py4j.commands.ConstructorCommand.invokeConstructor(ConstructorCommand.java:80)
    at py4j.commands.ConstructorCommand.execute(ConstructorCommand.java:69)
    at py4j.ClientServerConnection.waitForCommands(ClientServerConnection.java:182)
    at py4j.ClientServerConnection.run(ClientServerConnection.java:106)
    at java.lang.Thread.run(Thread.java:750)

从节点yarn-site.xml已配置:

<property>
<name>yarn.resourcemanager.hostname</name>
<value>hadoop-master</value>
</property>

解决步骤:

  1. 调整YARN超时与资源参数:在所有节点的yarn-site.xml中添加/修改以下配置,延长AM启动超时时间并匹配集群资源:
    <property>
        <name>yarn.resourcemanager.am.max-attempts</name>
        <value>3</value>
    </property>
    <property>
        <name>yarn.application-master.launcher.max-attempts</name>
        <value>3</value>
    </property>
    <property>
        <name>yarn.nodemanager.resource.memory-mb</name>
        <value>8192</value> <!-- 根据VM实际内存调整,确保大于executor内存请求 -->
    </property>
    <property>
        <name>yarn.scheduler.maximum-allocation-mb</name>
        <value>8192</value>
    </property>
    
    修改后重启YARN服务:stop-yarn.sh && start-yarn.sh
  2. 降低Spark资源请求:若VM内存不足,减少spark.executor.memory与spark.driver.memory的设置,避免资源申请失败:
    conf = SparkConf().setAppName("Spark_hadoop") \
                      .setMaster("yarn") \
                      .set("spark.executor.memory", "2g") \
                      .set("spark.driver.memory", "2g")
    
  3. 验证集群网络连通性:确保主从节点SSH无密码登录正常,所有节点可解析hadoop-master主机名(执行ping hadoop-master验证)。
  4. 查看YARN日志定位根因:在YARN仪表盘找到失败应用,点击ApplicationMaster的日志链接,查看具体启动失败细节(如依赖缺失、权限问题等)。

内容的提问来源于stack exchange,提问作者Karambit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 10:40:50