You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在AWS/EMR上配置Yarn/Spark+GeoSpark?容器启动异常求助

Fixing ExecutorLostFailure (Exit Status 52) on EMR Spark Cluster for GeoSpark Queries

Hey there, let's work through this issue step by step. You moved from a local cluster (facing GC overhead issues) to an EMR cluster with 1 master + 4 m4.xlarge nodes, but now you're hitting an ExecutorLostFailure with exit status 52. Let's break down what's going on and how to fix it:

First, Understand Exit Status 52

Exit code 52 in YARN typically signals memory-related container failures—even though you disabled physical/virtual memory checks, misaligned resource allocations or insufficient system resources can still trigger this. Let's start with the most obvious misconfigurations.

Step 1: Fix Executor Core & Memory Allocation (Critical!)

Your m4.xlarge nodes only have 4 vCPUs and 16GB RAM—but your current config asks for spark.executor.cores=5 and spark.driver.cores=5, which exceeds the total vCPUs per node. That's a guaranteed recipe for resource contention. Here's how to adjust:

  • Lower spark.executor.cores to 3 (leaves 1 core per node for YARN/NM and system processes)
  • Lower spark.driver.cores to 3 (same logic for the master node)
  • Reduce spark.executor.memory from 14.4g to 12g—the total executor memory + overhead (12g + 2g = 14g) leaves 2GB for the node's system processes (prevents the OS from killing containers due to memory starvation)
  • Adjust spark.yarn.executor.memoryOverhead to 2g (matches the 10-30% of executor memory recommended by Spark)

Step 2: Align YARN Node Resource Configs

Even with memory checks disabled, YARN needs to know the total resources available per node. Add these properties to your yarn-site classification:

{
  "Classification": "yarn-site",
  "Properties": {
    "yarn.nodemanager.vmem-check-enabled": "false",
    "yarn.nodemanager.pmem-check-enabled": "false",
    "yarn.nodemanager.resource.memory-mb": "16384", // 16GB in MB
    "yarn.scheduler.maximum-allocation-mb": "14336" // ~14GB, matches executor total
  }
}

Step 3: Optimize Spark Memory & GC (Prevent Repeat GC Issues)

You still need to address the original GC overhead problem while fixing the container failure:

  • Enable Kryo serialization (more efficient than Java serialization for GeoSpark spatial data):
    Add to spark-defaults:
    spark.serializer=org.apache.spark.serializer.KryoSerializer
    spark.kryo.registrator=org.datasyslab.geospark.serde.GeoSparkKryoRegistrator
    
  • Double-check spark.memory.fraction=0.7 and spark.memory.storageFraction=0.4—these are reasonable, but ensuring serialization is enabled will make this memory go further.
  • Adjust spark.default.parallelism to match your new core count: with 3 executors × 3 cores = 9 total cores, set this to 18 or 27 (2-3× core count is ideal for parallelism).

Step 4: Dig into Logs for Root Cause

To confirm exactly what's failing, check the detailed logs:

  • Use EMR Console to navigate to your application, then view the container logs for executor 2
  • SSH into the failing node (ip-172-31-16-70.eu-west-1.compute.internal) and check YARN NodeManager logs at /var/log/hadoop-yarn/yarn/yarn-nodemanager-*.log
  • Run yarn logs -applicationId <your-app-id> to pull all logs for the failed application—look for OOM errors, system memory pressure, or disk issues.

Step 5: Validate Executor Count

You have 4 worker nodes, so setting spark.executor.instances=3 is safe (leaves one worker node as a backup). If after adjusting resources you have headroom, you could bump this to 4, but start with 3 for stability.


内容的提问来源于stack exchange,提问作者Randomize

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:33:39