MapReduce多文件分析仅输出单条结果的实现问题
解决MapReduce全局最长单词长度输出问题
我看了你的代码,现在的问题是每个Mapper对应一个输入文件,都会输出自己文件里的最长单词长度,而Reducer现在会把每个Mapper的结果都输出一遍,所以才会有多行。要实现只输出全局最长的那一行,我们只需要调整Reducer的逻辑,让它只处理第一个(也就是最大的)键值对,然后就停止输出,同时修正代码里的小笔误就行。
问题分析
- Reducer逻辑问题:当前Reducer的
reduce方法会为每个Mapper输出的键(也就是每个文件的最长长度)执行一次,每次都写输出,所以会有N行(N是输入文件数)。而因为你已经设置了LongWritable.DecreasingComparator,Reducer收到的第一个键就是全局最大的长度,我们只需要输出这一个就行。 - 类名笔误:
main方法里的job.setJarByClass(WordCount.class);应该改成WordLength.class,因为你的主类名是WordLength,不然会报错找不到类。
修改后的完整代码
import java.io.IOException; import java.util.StringTokenizer; import org.apache.hadoop.conf.Configuration; import org.apache.hadoop.fs.Path; import org.apache.hadoop.io.IntWritable; import org.apache.hadoop.io.LongWritable; import org.apache.hadoop.io.Text; import org.apache.hadoop.mapreduce.Job; import org.apache.hadoop.mapreduce.Mapper; import org.apache.hadoop.mapreduce.Reducer; import org.apache.hadoop.mapreduce.lib.input.FileInputFormat; import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat; import org.apache.hadoop.util.GenericOptionsParser; public class WordLength { public static class Map extends Mapper<Object, Text, LongWritable, Text> { int max = Integer.MIN_VALUE; private Text word = new Text(); public void map(Object key, Text value, Context context) throws IOException, InterruptedException { String line = value.toString(); StringTokenizer tokenizer = new StringTokenizer(line); while (tokenizer.hasMoreTokens()) { String s = tokenizer.nextToken(); int val = s.length(); if (val > max) { max = val; word.set(s); } } } public void cleanup(Context context) throws IOException, InterruptedException { context.write(new LongWritable(max), word); } } public static class IntSumReducer extends Reducer<LongWritable, Text, Text, LongWritable> { // 添加一个标志位,控制只输出一次 private boolean hasOutput = false; public void reduce(LongWritable key, Iterable<Text> values, Context context) throws IOException, InterruptedException { // 只在第一次调用reduce时输出(也就是全局最大的那个值) if (!hasOutput) { context.write(new Text("longest"), key); hasOutput = true; } // 后续的reduce调用直接跳过,不输出 } } public static void main(String[] args) throws Exception { Configuration conf = new Configuration(); String[] otherArgs = new GenericOptionsParser(conf, args).getRemainingArgs(); if (otherArgs.length != 2) { System.err.println("Usage: wordlength <in> <out>"); System.exit(2); } Job job = Job.getInstance(conf, "word length"); // 修正类名笔误 job.setJarByClass(WordLength.class); job.setMapperClass(Map.class); job.setSortComparatorClass(LongWritable.DecreasingComparator.class); job.setNumReduceTasks(1); job.setReducerClass(IntSumReducer.class); job.setMapOutputKeyClass(LongWritable.class); job.setMapOutputValueClass(Text.class); job.setOutputKeyClass(Text.class); job.setOutputValueClass(LongWritable.class); FileInputFormat.addInputPath(job, new Path(otherArgs[0])); FileOutputFormat.setOutputPath(job, new Path(otherArgs[1])); System.exit(job.waitForCompletion(true) ? 0 : 1); } }
关键修改点说明
- Reducer添加标志位:新增
private boolean hasOutput = false;,在第一次reduce调用时输出结果,然后把hasOutput设为true,后续的reduce调用就不会再输出了,这样就只会保留全局最大的那一行。 - 修正主类名:把
job.setJarByClass(WordCount.class);改成job.setJarByClass(WordLength.class);,避免运行时找不到类的错误。 - 完善参数检查:添加了参数数量检查,如果输入输出路径不对,会提示用法,让程序更健壮。
这样修改后,不管你有多少个输入文件,最终只会输出一行结果,就是所有文件里最长单词的长度啦。
内容的提问来源于stack exchange,提问作者joy_jlee
相关产品推荐
相关产品推荐

