You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助解决Hadoop Yarn默认队列资源全为0致Spark作业无法运行

问题描述
  • 集群配置:3台服务器,每台2GB内存、2个CPU核心,已部署Hadoop 2.8.0集群,可正常运行Java MapReduce任务
  • 当前操作:搭建Spark 2.0.1-bin.without-hadoop + Scala环境,编写WordCount.scala并打包为JAR,执行spark-submit WordCount.jar /folder/file-1.txt提交作业
  • 异常现象:作业在YARN UI中始终处于ACCEPTED状态,无法启动

诊断信息

[Sun Nov 05 14:03:50 +0800 2023] Application is added to the scheduler and is not yet activated. Skipping AM assignment as cluster resource is empty. Details : AM Partition = <DEFAULT_PARTITION>; AM Resource Request = <memory:1024, vCores:1>; Queue Resource Limit for AM = <memory:0, vCores:0>; User AM Resource Limit of the queue = <memory:0, vCores:0>; Queue AM Resource Usage = <memory:0, vCores:0>;

队列状态页面显示:

  • Max Application Master Resources: <memory:0, vCores:0>
  • Used Application Master Resources: <memory:0, vCores:0>
  • Max Application Master Resources Per User: <memory:0, vCores:0>

当前capacity-scheduler.xml配置

<configuration>  
  <property>
    <name>yarn.scheduler.capacity.maximum-applications</name>
    <value>10000</value>
    <description>
      Maximum number of applications that can be pending and running.
    </description>
  </property>
  <property>
    <name>yarn.scheduler.capacity.maximum-am-resource-percent</name>
    <value>0.5</value>
    <description>
      Maximum percent of resources in the cluster which can be used to run 
      application masters i.e. controls number of concurrent running
      applications.
    </description>
  </property>
  <property>
    <name>yarn.scheduler.capacity.resource-calculator</name>
    <value>org.apache.hadoop.yarn.util.resource.DefaultResourceCalculator</value>
    <description>
      The ResourceCalculator implementation to be used to compare 
      Resources in the scheduler.
      The default i.e. DefaultResourceCalculator only uses Memory while
      DominantResourceCalculator uses dominant-resource to compare 
      multi-dimensional resources such as Memory, CPU etc.
    </description>
  </property>
  <property>
    <name>yarn.scheduler.capacity.root.queues</name>
    <value>default</value>
    <description>
      The queues at the this level (root is the root queue).
    </description>
  </property>
  <property>
    <name>yarn.scheduler.capacity.root.default.capacity</name>
    <value>100</value>
    <description>Default queue target capacity.</description>
  </property>
  <property>
    <name>yarn.scheduler.capacity.root.default.user-limit-factor</name>
    <value>1</value>
    <description>
      Default queue user limit a percentage from 0.0 to 1.0.
    </description>
  </property>
  <property>
    <name>yarn.scheduler.capacity.root.default.maximum-capacity</name>
    <value>100</value>
    <description>
      The maximum capacity of the default queue. 
    </description>
  </property>
  <!-- Ensure the default queue has minimum capacity -->
  <property>
    <name>yarn.scheduler.capacity.root.default.minimum-capacity</name>
    <value>1</value> <!-- This ensures that the default queue has some minimum capacity -->
    <description>
      The minimum capacity of the default queue.
    </description>
  </property>
  <property>
    <name>yarn.scheduler.capacity.root.default.state</name>
    <value>RUNNING</value>
    <description>
      The state of the default queue. State can be one of RUNNING or STOPPED.
    </description>
  </property>
  <property>
    <name>yarn.scheduler.capacity.root.default.acl_submit_applications</name>
    <value>*</value>
    <description>
      The ACL of who can submit jobs to the default queue.
    </description>
  </property>
  <property>
    <name>yarn.scheduler.capacity.root.default.acl_administer_queue</name>
    <value>*</value>
    <description>
      The ACL of who can administer jobs on the default queue.
    </description>
  </property>
  <property>
    <name>yarn.scheduler.capacity.node-locality-delay</name>
    <value>40</value>
    <description>
      Number of missed scheduling opportunities after which the CapacityScheduler 
      attempts to schedule rack-local containers. 
      Typically this should be set to number of nodes in the cluster, By default is setting 
      approximately number of nodes in one rack which is 40.
    </description>
  </property>
  <property>
    <name>yarn.scheduler.capacity.queue-mappings</name>
    <value></value>
    <description>
      A list of mappings that will be used to assign jobs to queues
      The syntax for this list is [u|g]:[name]:[queue_name][,next mapping]*
      Typically this list will be used to map users to queues,
      for example, u:%user:%user maps all users to queues with the same name
      as the user.
    </description>
  </property>
  <property>
    <name>yarn.scheduler.capacity.queue-mappings-override.enable</name>
    <value>false</value>
    <description>
      If a queue mapping is present, will it override the value specified
      by the user? This can be used by administrators to place jobs in queues
      that are different than the one specified by the user.
      The default is false.
    </description>
  </property>
  <property>
    <name>yarn.scheduler.capacity.per-node-heartbeat.maximum-offswitch-assignments</name>
    <value>1</value>
    <description>
      Controls the number of OFF_SWITCH assignments allowed
      during a node's heartbeat. Increasing this value can improve
      scheduling rate for OFF_SWITCH containers. Lower values reduce
      "clumping" of applications on particular nodes. The default is 1.
      Legal values are 1-MAX_INT. This config is refreshable.
    </description>
  </property>
  <property>
          <name>yarn.scheduler.capacity.root.maximum-allocation-mb</name>
          <value>1536</value>
  </property>
    <property>
          <name>yarn.scheduler.capacity.root.maximum-allocation-vcores</name>
          <value>2</value>
  </property>
  <property>
          <name>yarn.scheduler.maximum-allocation-mb</name>
          <value>1536</value>
  </property>
  <property>
          <name>yarn.scheduler.maximum-allocation-vcores</name>
          <value>2</value>
  </property>
  <property>
          <name>yarn.nodemanager.resource.cpu-vcores </name>
          <value>6</value>
  </property>
  <property>
          <name>yarn.nodemanager.resource.memory-mb </name>
          <value>1536</value>
  </property>
</configuration>

问题原因及修复方案

核心原因

  1. CPU资源配置不匹配:每台服务器实际只有2个CPU核心,但配置中yarn.nodemanager.resource.cpu-vcores设为6,YARN无法识别虚假的CPU核心数,导致集群总可用vCores计算异常
  2. 资源计算器未启用多维资源调度:当前使用DefaultResourceCalculator仅计算内存,而Spark作业同时请求内存和CPU,导致队列AM资源限制计算错误

修复步骤

1. 修正NodeManager CPU核心配置

将yarn.nodemanager.resource.cpu-vcores的值改为服务器实际的CPU核心数2:

<property>
  <name>yarn.nodemanager.resource.cpu-vcores</name>
  <value>2</value>
</property>

2. 启用多维资源计算器

将资源计算器改为DominantResourceCalculator,支持同时计算内存和CPU资源:

<property>
  <name>yarn.scheduler.capacity.resource-calculator</name>
  <value>org.apache.hadoop.yarn.util.resource.DominantResourceCalculator</value>
</property>

3. 验证队列AM资源配置

yarn.scheduler.capacity.maximum-am-resource-percent保持0.5即可(允许集群50%资源用于AM),该配置本身无问题,修正前因资源计算错误导致队列AM资源限制为0

4. 重启Hadoop集群

修改配置后,依次重启YARN服务:

# 停止YARN服务
stop-yarn.sh
# 启动YARN服务
start-yarn.sh

5. 验证修复效果

提交Spark作业后,检查YARN UI的队列状态,确认Max Application Master Resources显示正常数值(例如内存约2GB,vCores约3),作业能正常从ACCEPTED转为RUNNING

内容的提问来源于stack exchange,提问作者Zifeng Tan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 23:07:07