You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS EMR启动遇InvalidAuxServiceException,如何预配置mapreduce_shuffle?

Fixing "mapreduce_shuffle does not exist" Error in AWS EMR (Cluster Launch Time Configuration)

Got it, let's tackle this EMR cluster issue you're hitting. That org.apache.hadoop.yarn.exceptions.InvalidAuxServiceException error happens because YARN isn't configured to recognize the mapreduce_shuffle auxiliary service by default in your setup—and you're right, manually modifying nodes post-launch is a hassle. The good news is you can bake this configuration directly into your cluster launch process so it's applied automatically from the start.

Here are the step-by-step ways to do this:

1. Using the AWS EMR Console

When creating your cluster via the console:

  • Navigate to the Software configuration section
  • Click Edit software settings
  • Select the Enter configuration option, then paste this JSON snippet:
    [
      {
        "Classification": "yarn-site",
        "Properties": {
          "yarn.nodemanager.aux-services": "mapreduce_shuffle",
          "yarn.nodemanager.aux-services.mapreduce_shuffle.class": "org.apache.hadoop.mapred.ShuffleHandler"
        }
      }
    ]
    
  • Save the settings and proceed with cluster creation.

This tells EMR to inject these properties into the yarn-site.xml file on all relevant nodes (Master, Core, Task) during launch. The first property declares the auxiliary service, and the second points to its implementation class.

2. Using the AWS CLI

If you prefer launching clusters via the CLI, create a JSON config file (e.g., yarn-shuffle-config.json) with the same content as above, then reference it in your create-cluster command:

aws emr create-cluster \
  --name "Auto-Terminate Task Cluster" \
  --release-label emr-6.15.0 \
  --instance-type m5.xlarge \
  --instance-count 3 \
  --configurations file://yarn-shuffle-config.json \
  --steps Type=Spark,Name="My Job",ActionOnFailure=TERMINATE_CLUSTER,Args=[--class,com.example.MyTask,s3://my-bucket/jars/my-task.jar] \
  --auto-terminate

Replace the release label, instance details, and step arguments with your specific values.

3. Using CloudFormation (Infrastructure as Code)

If you use CloudFormation to provision EMR clusters, add the Configurations block to your cluster resource definition:

Resources:
  MyEMRTaskCluster:
    Type: AWS::EMR::Cluster
    Properties:
      Name: "Auto-Terminate Task Cluster"
      ReleaseLabel: emr-6.15.0
      Instances:
        InstanceGroups:
          - InstanceCount: 1
            InstanceRole: MASTER
            InstanceType: m5.xlarge
          - InstanceCount: 2
            InstanceRole: CORE
            InstanceType: m5.xlarge
        AutoTerminate: true
      Configurations:
        - Classification: "yarn-site"
          Properties:
            yarn.nodemanager.aux-services: "mapreduce_shuffle"
            yarn.nodemanager.aux-services.mapreduce_shuffle.class: "org.apache.hadoop.mapred.ShuffleHandler"
      Steps:
        - Name: "Run My Job"
          ActionOnFailure: TERMINATE_CLUSTER
          HadoopJarStep:
            Jar: "command-runner.jar"
            Args: ["spark-submit", "--class", "com.example.MyTask", "s3://my-bucket/jars/my-task.jar"]

Why This Works

EMR's configuration classification system automatically propagates these settings to all nodes' yarn-site.xml files during cluster initialization. You won't need to manually edit files or restart YARN services—everything is set up before your job starts running.

Just make sure the EMR release label you're using supports these properties (all modern EMR versions from 5.x onwards do, so you shouldn't run into issues).

内容的提问来源于stack exchange,提问作者Sateesh K

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 10:04:50