2018年能否在Amazon EMR中设置HADOOP_HEAPSIZE解决多文件OOM问题?
Absolutely! By 2018, Amazon EMR fully supports configuring HADOOP_HEAPSIZE (along with more granular heap settings for MapReduce tasks) to tackle OOM errors caused by millions of input files. You have two straightforward options—using the C# EMR API directly, or leveraging Bootstrap scripts. Let’s break them down:
You can set heap size parameters directly when creating your EMR cluster or submitting jobs using the AWS SDK for .NET. This is the cleanest approach if you want to embed the configuration in your code.
Example Code for Cluster Creation:
using Amazon.ElasticMapReduce; using Amazon.ElasticMapReduce.Model; var emrClient = new AmazonElasticMapReduceClient(); var createClusterRequest = new RunJobFlowRequest { Name = "Heap-Optimized Cluster for Large Inputs", ReleaseLabel = "emr-5.15.0", // A 2018-compatible EMR release that supports these configs Instances = new JobFlowInstancesConfig { // Add your instance type, count, subnet, etc. here MasterInstanceType = "m4.xlarge", SlaveInstanceType = "m4.xlarge", InstanceCount = 3 }, Configurations = new List<Configuration> { // Configure global Hadoop daemon heap size new Configuration { Classification = "hadoop-env", Properties = new Dictionary<string, string> { { "HADOOP_HEAPSIZE", "4096" } // Sets heap size to 4GB (adjust based on your instance size) } }, // Configure MapReduce task-specific JVM heap sizes (critical for job-level OOM fixes) new Configuration { Classification = "mapred-site", Properties = new Dictionary<string, string> { { "mapreduce.map.memory.mb", "4096" }, // Allocate 4GB container for Map tasks { "mapreduce.reduce.memory.mb", "8192" }, // Allocate 8GB container for Reduce tasks { "mapreduce.map.java.opts", "-Xmx3072m" }, // Set heap to ~75% of container size (best practice) { "mapreduce.reduce.java.opts", "-Xmx6144m" } } } }, // Add your job steps, service roles, etc. here }; var clusterResponse = emrClient.RunJobFlow(createClusterRequest);
Key Notes:
HADOOP_HEAPSIZEin thehadoop-envclassification adjusts heap for Hadoop daemons (like NameNode, DataNode). For job-specific OOMs, focus on themapred-siteparameters—they control the JVM heap for your Map/Reduce tasks, which is usually the culprit when processing millions of small files.- Always align the JVM heap size (
-Xmx) to ~75% of the container memory (mapreduce.map.memory.mb) to leave room for system overhead.
If you need more flexibility (e.g., applying different settings to master vs. core nodes), a Bootstrap script is a great option. Here's how to set it up:
Step 1: Create the Bootstrap Script
Save this as set-heap-size.sh and upload it to an S3 bucket:
#!/bin/bash # Set global Hadoop daemon heap size echo "export HADOOP_HEAPSIZE=4096" >> /etc/hadoop/conf/hadoop-env.sh # Set MapReduce task heap sizes cat >> /etc/hadoop/conf/mapred-site.xml << EOF <property> <name>mapreduce.map.java.opts</name> <value>-Xmx3072m</value> </property> <property> <name>mapreduce.reduce.java.opts</name> <value>-Xmx6144m</value> </property> EOF # Restart Hadoop services to apply changes (optional for some EMR versions) sudo stop hadoop-datanode sudo start hadoop-datanode sudo stop hadoop-namenode sudo start hadoop-namenode
Step 2: Reference the Script in C# API
var createClusterRequest = new RunJobFlowRequest { // ... existing cluster configuration ... BootstrapActions = new List<BootstrapActionConfig> { new BootstrapActionConfig { Name = "Configure Hadoop Heap Sizes", ScriptBootstrapAction = new ScriptBootstrapActionConfig { Path = "s3://your-bucket-name/path/to/set-heap-size.sh", Args = new List<string>() // Add arguments if your script accepts them } } } };
Adjusting heap size helps, but combining it with these fixes will make your job far more stable:
- Merge Small Files: Use
CombineFileInputFormatin your MapReduce job to group small files into larger splits, reducing the number of Map tasks (and thus memory overhead from task tracking). - Adjust Split Size: Set
mapreduce.input.fileinputformat.split.maxsizeinmapred-siteto increase the maximum split size, which reduces the total number of Map tasks. - Enable JVM Reuse: Set
mapreduce.job.jvm.numtasksto a value >1 to reuse JVMs across tasks, cutting down on memory allocation overhead.
内容的提问来源于stack exchange,提问作者user2330278

