MapReduce Reduce Task中Iterable失效求助——Hadoop WordCount程序异常
Fixing Your Hadoop WordCount Reducer Issue
Hey there! Let's get your WordCount program working as expected. The problem you're facing—where "hadoop" shows up twice with a count of 1 instead of once with 2—almost always boils down to how your reducer is handling input values.
Common Root Cause
Chances are, your reducer is writing out each individual value for the key instead of summing all values together first. For example, if you’re looping through the values iterator and calling context.write() inside the loop, you’ll get a separate entry for every count emitted by the mapper.
Corrected Code Example
Here’s a complete, working version of the WordCount program with the fix in the reducer:
Mapper Class
package in.edureka.mapreduce; import org.apache.hadoop.io.IntWritable; import org.apache.hadoop.io.LongWritable; import org.apache.hadoop.io.Text; import org.apache.hadoop.mapreduce.Mapper; import java.io.IOException; public class WordCountMapper extends Mapper<LongWritable, Text, Text, IntWritable> { private final static IntWritable one = new IntWritable(1); private Text word = new Text(); @Override protected void map(LongWritable key, Text value, Context context) throws IOException, InterruptedException { // Split input line into words String[] words = value.toString().split("\\s+"); for (String w : words) { word.set(w.toLowerCase()); // Optional: normalize to lowercase to avoid case sensitivity context.write(word, one); } } }
Reducer Class (The Fix Is Here!)
package in.edureka.mapreduce; import org.apache.hadoop.io.IntWritable; import org.apache.hadoop.io.Text; import org.apache.hadoop.mapreduce.Reducer; import java.io.IOException; public class WordCountReducer extends Reducer<Text, IntWritable, Text, IntWritable> { @Override protected void reduce(Text key, Iterable<IntWritable> values, Context context) throws IOException, InterruptedException { int sum = 0; // Iterate through all values for the current key and sum them up for (IntWritable val : values) { sum += val.get(); } // Write the key and the total sum ONCE, not per value context.write(key, new IntWritable(sum)); } }
Driver Class
package in.edureka.mapreduce; import org.apache.hadoop.conf.Configuration; import org.apache.hadoop.fs.Path; import org.apache.hadoop.io.IntWritable; import org.apache.hadoop.io.Text; import org.apache.hadoop.mapreduce.Job; import org.apache.hadoop.mapreduce.lib.input.FileInputFormat; import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat; import java.io.IOException; public class WordCountDriver { public static void main(String[] args) throws IOException, InterruptedException, ClassNotFoundException { Configuration conf = new Configuration(); Job job = Job.getInstance(conf, "WordCount"); job.setJarByClass(WordCountDriver.class); job.setMapperClass(WordCountMapper.class); job.setCombinerClass(WordCountReducer.class); // Optional: reduces network traffic by aggregating at mapper nodes job.setReducerClass(WordCountReducer.class); job.setOutputKeyClass(Text.class); job.setOutputValueClass(IntWritable.class); FileInputFormat.addInputPath(job, new Path(args[0])); FileOutputFormat.setOutputPath(job, new Path(args[1])); System.exit(job.waitForCompletion(true) ? 0 : 1); } }
Key Fixes Explained
- Reducer Sum Logic: We initialize a
sumvariable to 0, loop through all values for the key, and add each value to the sum. Only after processing all values do we write the key and total sum to the output. - Optional Combiner: Reusing the reducer logic as a combiner reduces the amount of data sent to the reducer by aggregating counts at the mapper node first—this boosts performance for large datasets.
- Case Normalization: The mapper converts words to lowercase so "Hadoop" and "hadoop" are treated as the same key (remove this if you need case-sensitive counting).
How to Test
- Compile and package the code into a JAR file.
- Run it with your input data using the Hadoop command:
hadoop jar your-wordcount.jar in.edureka.mapreduce.WordCountDriver /path/to/input /path/to/output - Check the output file—you should now see
this 1andhadoop 2as expected.
内容的提问来源于stack exchange,提问作者Qasim
相关产品推荐
相关产品推荐

