You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MapReduce Reduce Task中Iterable失效求助——Hadoop WordCount程序异常

Fixing Your Hadoop WordCount Reducer Issue

Hey there! Let's get your WordCount program working as expected. The problem you're facing—where "hadoop" shows up twice with a count of 1 instead of once with 2—almost always boils down to how your reducer is handling input values.

Common Root Cause

Chances are, your reducer is writing out each individual value for the key instead of summing all values together first. For example, if you’re looping through the values iterator and calling context.write() inside the loop, you’ll get a separate entry for every count emitted by the mapper.

Corrected Code Example

Here’s a complete, working version of the WordCount program with the fix in the reducer:

Mapper Class

package in.edureka.mapreduce;

import org.apache.hadoop.io.IntWritable;
import org.apache.hadoop.io.LongWritable;
import org.apache.hadoop.io.Text;
import org.apache.hadoop.mapreduce.Mapper;

import java.io.IOException;

public class WordCountMapper extends Mapper<LongWritable, Text, Text, IntWritable> {
    private final static IntWritable one = new IntWritable(1);
    private Text word = new Text();

    @Override
    protected void map(LongWritable key, Text value, Context context) throws IOException, InterruptedException {
        // Split input line into words
        String[] words = value.toString().split("\\s+");
        for (String w : words) {
            word.set(w.toLowerCase()); // Optional: normalize to lowercase to avoid case sensitivity
            context.write(word, one);
        }
    }
}

Reducer Class (The Fix Is Here!)

package in.edureka.mapreduce;

import org.apache.hadoop.io.IntWritable;
import org.apache.hadoop.io.Text;
import org.apache.hadoop.mapreduce.Reducer;

import java.io.IOException;

public class WordCountReducer extends Reducer<Text, IntWritable, Text, IntWritable> {
    @Override
    protected void reduce(Text key, Iterable<IntWritable> values, Context context) throws IOException, InterruptedException {
        int sum = 0;
        // Iterate through all values for the current key and sum them up
        for (IntWritable val : values) {
            sum += val.get();
        }
        // Write the key and the total sum ONCE, not per value
        context.write(key, new IntWritable(sum));
    }
}

Driver Class

package in.edureka.mapreduce;

import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.fs.Path;
import org.apache.hadoop.io.IntWritable;
import org.apache.hadoop.io.Text;
import org.apache.hadoop.mapreduce.Job;
import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;
import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;

import java.io.IOException;

public class WordCountDriver {
    public static void main(String[] args) throws IOException, InterruptedException, ClassNotFoundException {
        Configuration conf = new Configuration();
        Job job = Job.getInstance(conf, "WordCount");
        
        job.setJarByClass(WordCountDriver.class);
        job.setMapperClass(WordCountMapper.class);
        job.setCombinerClass(WordCountReducer.class); // Optional: reduces network traffic by aggregating at mapper nodes
        job.setReducerClass(WordCountReducer.class);
        
        job.setOutputKeyClass(Text.class);
        job.setOutputValueClass(IntWritable.class);
        
        FileInputFormat.addInputPath(job, new Path(args[0]));
        FileOutputFormat.setOutputPath(job, new Path(args[1]));
        
        System.exit(job.waitForCompletion(true) ? 0 : 1);
    }
}

Key Fixes Explained

  • Reducer Sum Logic: We initialize a sum variable to 0, loop through all values for the key, and add each value to the sum. Only after processing all values do we write the key and total sum to the output.
  • Optional Combiner: Reusing the reducer logic as a combiner reduces the amount of data sent to the reducer by aggregating counts at the mapper node first—this boosts performance for large datasets.
  • Case Normalization: The mapper converts words to lowercase so "Hadoop" and "hadoop" are treated as the same key (remove this if you need case-sensitive counting).

How to Test

  1. Compile and package the code into a JAR file.
  2. Run it with your input data using the Hadoop command:
    hadoop jar your-wordcount.jar in.edureka.mapreduce.WordCountDriver /path/to/input /path/to/output
    
  3. Check the output file—you should now see this 1 and hadoop 2 as expected.

内容的提问来源于stack exchange,提问作者Qasim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:34:09