如何在AWS/EMR上配置Yarn/Spark+GeoSpark?容器启动异常求助
Hey there, let's work through this issue step by step. You moved from a local cluster (facing GC overhead issues) to an EMR cluster with 1 master + 4 m4.xlarge nodes, but now you're hitting an ExecutorLostFailure with exit status 52. Let's break down what's going on and how to fix it:
First, Understand Exit Status 52
Exit code 52 in YARN typically signals memory-related container failures—even though you disabled physical/virtual memory checks, misaligned resource allocations or insufficient system resources can still trigger this. Let's start with the most obvious misconfigurations.
Step 1: Fix Executor Core & Memory Allocation (Critical!)
Your m4.xlarge nodes only have 4 vCPUs and 16GB RAM—but your current config asks for spark.executor.cores=5 and spark.driver.cores=5, which exceeds the total vCPUs per node. That's a guaranteed recipe for resource contention. Here's how to adjust:
- Lower
spark.executor.coresto3(leaves 1 core per node for YARN/NM and system processes) - Lower
spark.driver.coresto3(same logic for the master node) - Reduce
spark.executor.memoryfrom14.4gto12g—the total executor memory + overhead (12g + 2g = 14g) leaves 2GB for the node's system processes (prevents the OS from killing containers due to memory starvation) - Adjust
spark.yarn.executor.memoryOverheadto2g(matches the 10-30% of executor memory recommended by Spark)
Step 2: Align YARN Node Resource Configs
Even with memory checks disabled, YARN needs to know the total resources available per node. Add these properties to your yarn-site classification:
{ "Classification": "yarn-site", "Properties": { "yarn.nodemanager.vmem-check-enabled": "false", "yarn.nodemanager.pmem-check-enabled": "false", "yarn.nodemanager.resource.memory-mb": "16384", // 16GB in MB "yarn.scheduler.maximum-allocation-mb": "14336" // ~14GB, matches executor total } }
Step 3: Optimize Spark Memory & GC (Prevent Repeat GC Issues)
You still need to address the original GC overhead problem while fixing the container failure:
- Enable Kryo serialization (more efficient than Java serialization for GeoSpark spatial data):
Add tospark-defaults:spark.serializer=org.apache.spark.serializer.KryoSerializer spark.kryo.registrator=org.datasyslab.geospark.serde.GeoSparkKryoRegistrator - Double-check
spark.memory.fraction=0.7andspark.memory.storageFraction=0.4—these are reasonable, but ensuring serialization is enabled will make this memory go further. - Adjust
spark.default.parallelismto match your new core count: with 3 executors × 3 cores = 9 total cores, set this to18or27(2-3× core count is ideal for parallelism).
Step 4: Dig into Logs for Root Cause
To confirm exactly what's failing, check the detailed logs:
- Use EMR Console to navigate to your application, then view the container logs for executor 2
- SSH into the failing node (
ip-172-31-16-70.eu-west-1.compute.internal) and check YARN NodeManager logs at/var/log/hadoop-yarn/yarn/yarn-nodemanager-*.log - Run
yarn logs -applicationId <your-app-id>to pull all logs for the failed application—look for OOM errors, system memory pressure, or disk issues.
Step 5: Validate Executor Count
You have 4 worker nodes, so setting spark.executor.instances=3 is safe (leaves one worker node as a backup). If after adjusting resources you have headroom, you could bump this to 4, but start with 3 for stability.
内容的提问来源于stack exchange,提问作者Randomize

