You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Hadoop MapReduce Reduce方法未调用及输出格式问题咨询

问题分析与解决思路

一、Reducer方法不被调用+@Override报错的核心原因

你的Reducer类泛型参数和reduce方法参数都出现了关键错误,这直接导致了注解报错和方法不被调用:

  1. Reducer泛型定义完全错误
    你当前的Reducer泛型写的是:
public class PopularityReducer extends Reducer<Text, Iterable<String>, Text, Text>

Reducer的泛型规则是 <KEYIN, VALUEIN, KEYOUT, VALUEOUT>,其中:

  • KEYIN:Mapper输出的Key类型,这里你的Mapper输出Key是Text,没问题
  • VALUEIN:Mapper输出的Value类型,你的Mapper输出Value是Text,但你错误地写成了Iterable<String>——Iterable是reduce方法中接收values集合的类型,根本不是泛型里的VALUEIN参数。

正确的泛型定义应该是:

public class PopularityReducer extends Reducer<Text, Text, Text, Text>
  1. reduce方法参数类型不匹配
    对应的reduce方法里的Iterable<String> values也不对,Hadoop传递数据用的是自身序列化类型(实现Writable接口的类),这里应该用Iterable<Text>,遍历的时候再把Text转成String:
@Override
protected void reduce(Text key, Iterable<Text> values, Context context) throws IOException, InterruptedException {
    int sum = 0;
    for (Text val : values) {
        String preference = val.toString();
        // 注意:字符串比较必须用equals(),不能用==,否则会因为对象引用不同判断失效
        if ("true".equals(preference)) {
            sum += 1;
        } else if ("false".equals(preference)) {
            sum -= 1;
        }
    }
    context.write(key, new Text(String.valueOf(sum)));
}
  1. Mapper的致命错误
    你的Mapper里直接读取本地文件src\\testinput.json,这完全不符合MapReduce的设计逻辑!Mapper应该处理FileInputFormat指定的输入路径中的数据,要通过value参数获取输入内容,而不是自己读本地文件。这种写法在本地模式下可能凑合用,但提交到集群后,每个Task节点都没有这个本地文件,直接就会报错。

修改后的Mapper应该是这样:

@Override
protected void map(Text key, Text value, Context context) throws IOException, InterruptedException {
    JSONParser jsonParser = new JSONParser();
    try {
        // 解析value传递的输入文件内容(如果你的JSON是单行格式)
        JSONObject jsonobject = (JSONObject) jsonParser.parse(value.toString());
        JSONArray jsonArray = (JSONArray) jsonobject.get("votes");
        Iterator<JSONObject> iterator = jsonArray.iterator();
        while(iterator.hasNext()) {
            JSONObject obj = iterator.next();
            String song_id_rave_id = (String) obj.get("song_ID") + "|" + (String) obj.get("rave_ID");
            String preference = (String) obj.get("preference");
            System.out.println(song_id_rave_id + "||" + preference);
            context.write(new Text(song_id_rave_id), new Text(preference));
        }
    } catch(ParseException e) {
        e.printStackTrace();
    }
}

二、关于输出为.txt文件的问题

你已经在代码中设置了job.setOutputFormatClass(TextOutputFormat.class);,TextOutputFormat就是Hadoop默认的文本输出格式——它生成的part-r-00000这类文件本质就是纯文本格式,你可以直接下载后重命名为.txt,或者在读取时直接当作txt文件处理,完全满足你的需求。如果一定要自定义文件名,可以实现自定义OutputFormat,但一般默认的格式就够用了。

其他需要修正的小细节

  1. 在run方法中,你设置了job.setOutputValueClass(IntWritable.class);,但你的Reducer输出的Value是Text类型,这里要改成job.setOutputValueClass(Text.class);,否则会出现类型不匹配错误。
  2. 你的输入JSON最后一个对象后面多了个多余的逗号:"rave_ID": "rave001", },,这会导致JSON解析失败,修正后的JSON:
{"votes":[
    { "song_ID": "Piece of your heart", "mbr_ID": "001", "preference": "true", "timestamp": "11:22:33", "rave_ID": "rave001" }, 
    { "song_ID": "Piece of your heart", "mbr_ID": "002", "preference": "true", "timestamp": "11:22:33", "rave_ID": "rave001" }, 
    { "song_ID": "Atje voor de sfeer", "mbr_ID": "001", "preference": "false", "timestamp": "11:44:33", "rave_ID": "rave001" }, 
    { "song_ID": "Atje voor de sfeer", "mbr_ID": "002", "preference": "false", "timestamp": "11:44:33", "rave_ID": "rave001" }, 
    { "song_ID": "Atje voor de sfeer", "mbr_ID": "003", "preference": "true", "timestamp": "11:44:33", "rave_ID": "rave001" }
]}

内容的提问来源于stack exchange,提问作者TNelen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:20:17