EMR集群运行时修改yarn-site.xml配置无效及内存超限问题求助
Hey there, let's work through this EMR memory problem step by step. First, we'll clarify why your manual yarn-site.xml tweak didn't work, then cover the proper ways to update configurations on a running cluster, plus extra tips to address the root memory issue.
Why modifying yarn-site.xml directly and restarting RM didn't work
EMR manages configurations through its centralized configuration system—not just local XML files. When you edit yarn-site.xml manually, EMR's config service will overwrite your changes on the next sync. Also, restarting only the Resource Manager isn't enough: Node Managers need to pick up the new config too, and the EMR config UI only displays settings applied through official channels.
How to update YARN configs on a running EMR cluster
You have two reliable, official methods to adjust settings without recreating the cluster:
1. Use the AWS CLI
Run this command to push the config change to your cluster. EMR will automatically sync it to all nodes and restart the required services:
aws emr modify-cluster-attributes --cluster-id <your-cluster-id> --configuration '[{"Classification":"yarn-site","Properties":{"yarn.nodemanager.vmem-check-enabled":"false"}}]'
Wait 5-10 minutes for EMR to complete the sync and service restarts. The config will then show up in the EMR config page (namenode:8088/conf).
2. Use the AWS EMR Console
- Navigate to the EMR service in the AWS Console, locate your cluster
- Click Actions → Edit cluster configurations
- Add a new configuration with the
yarn-siteclassification, then add the propertyyarn.nodemanager.vmem-check-enabledset tofalse - Save the changes—EMR handles syncing to nodes and restarting services automatically
Is yarn.nodemanager.vmem-check-enabled only configurable at cluster creation?
Absolutely not! Both methods above let you modify this setting on a running cluster. You don't need to spin up a new cluster just to adjust this config.
Bonus tips to fix the root memory issue
Disabling virtual memory checks is a quick band-aid—addressing the actual memory overuse is better:
- Verify job resource settings: Make sure your job's memory requests (like
mapreduce.map.memory.mb/mapreduce.reduce.memory.mbfor MapReduce, orspark.executor.memoryfor Spark) are reasonable and don't exceed the node's available physical memory. - Adjust YARN node memory allocation: Tweak
yarn.nodemanager.resource.memory-mbto properly partition your node's memory across containers, preventing any single container from hogging too much. - Monitor memory usage: Use CloudWatch metrics (EMR provides built-in ones) or the YARN UI to track node memory utilization, then optimize your job's resource requests accordingly.
Pro tip: After applying config changes, re-submit your job and check the YARN UI to confirm the memory issue is resolved.
内容的提问来源于stack exchange,提问作者Marcos Santana

