You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行MapReduce专利统计应用时出现值类型不匹配错误该如何解决

MapReduce代码错误修复方案

报错核心原因

你在run方法中重复设置了作业的输出值类型,后设置的FloatWritable覆盖了先设置的IntWritable,而Hadoop默认将作业全局的输出值类型作为Mapper的输出值类型校验标准,你的Mapper实际输出的是IntWritable,因此触发类型不匹配报错。

具体修复步骤

  • 修正类型配置

删除原来重复的job.setOutputValueClass配置,分开指定Mapper和Reducer的输出类型,在run方法中替换原有类型配置代码为:

// 指定Mapper输出类型
job.setMapOutputKeyClass(Text.class);
job.setMapOutputValueClass(IntWritable.class);
// 指定最终Reducer输出类型
job.setOutputKeyClass(Text.class);
job.setOutputValueClass(FloatWritable.class);
// 必须设置Reduce任务数为2,匹配Partitioner的分区规则
job.setNumReduceTasks(2);
  • 移除错误的Combiner配置

你直接将输出为FloatWritable的Reducer设为Combiner,但Combiner的输出类型必须和Mapper的输出类型(IntWritable)一致,且求平均值的场景本身不适合用Combiner(局部平均求和后无法直接合并为全局平均,会导致统计结果错误),直接删除以下代码:

job.setCombinerClass(MyReducer.class);
  • 修正Partitioner逻辑

你将国家代码转成小写后匹配大写的AU,永远无法匹配成功,同时要兼容Reduce任务数不为2的场景避免分区越界,修改MyPartitioner代码为:

public static class MyPartitioner extends Partitioner<Text, IntWritable> {
    public int getPartition(Text key, IntWritable value, int numReduceTasks) {
        if(numReduceTasks <= 1) {
            return 0;
        }
        String country = key.toString();
        if ("AU".equals(country)) {
            return 0;
        } else {
            return 1;
        }
    }
}
  • 修正Mapper逻辑漏洞

现有代码会把1975年之前的专利也输出,且没有处理CLAIMS字段为空、格式异常的脏数据,修改map方法为:

@Override
public void map(Object key, Text value, Context context) throws IOException, InterruptedException {
    String line = value.toString();
    // 跳过表头
    if (line.contains("PATENT")) {
        return; 
    }
    String[] patentData = line.split(",");
    try {
        int gyear = Integer.parseInt(patentData[1]);
        String claimsStr = patentData[8].trim();
        // 仅统计1975年之后、CLAIMS字段有效的专利
        if (gyear >= 1975 && !claimsStr.isEmpty()) {
            int claims = Integer.parseInt(claimsStr);
            context.write(new Text(patentData[4]), new IntWritable(claims));
        }
    } catch (NumberFormatException e) {
        // 格式异常的脏数据直接跳过
        return;
    }
}

结果查看方式

作业运行成功后,输出目录下的part-r-00000文件就是AU国家的统计结果,直接执行以下命令查看即可:
hadoop fs -cat 你的输出目录路径/part-r-00000

内容的提问来源于stack exchange,提问作者Thi Nguyen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 10:57:01