为何Hadoop中Map类无输出记录?输入记录数达100000
问题排查:Map输出记录数为0的原因
1. Mapper输出类型与Job配置不匹配
你的Mapper类定义的输出键类型是IntWritable,值类型是Text:
public static class MapClass extends Mapper<LongWritable,Text,IntWritable,Text>
但Driver中设置的全局输出键值类型是Text和LongWritable:
job.setOutputKeyClass(Text.class); job.setOutputValueClass(LongWritable.class);
MapReduce框架会严格校验Mapper输出类型与Job配置的一致性,不匹配时会直接丢弃所有Map输出记录,这是导致输出为0的核心原因。
2. 异常信息无法有效查看
Mapper的catch块中使用System.out.println(e.getMessage())打印异常,但在MapReduce任务中,System.out的输出不会被收集到任务日志中,导致你无法定位是否存在解析错误(比如数据包含表头行、某行数据格式异常导致数组越界或转Integer失败)。
3. 潜在的数据格式问题
- 若原始数据包含表头行(第一行是列名),解析第一行时
str[1]会是字符串(比如"Year"),调用Integer.parseInt会抛出异常,进入catch块后无输出。 - 部分行可能存在格式错误(比如缺失列、逗号分隔异常),导致
str数组长度不足2,访问str[1]抛出ArrayIndexOutOfBoundsException。
修复方案
修正Job输出类型配置:
将Driver中的输出类型改为与Mapper一致,若Reducer输出类型不同,需单独设置Map阶段的输出类型:// 匹配Mapper的输出类型 job.setMapOutputKeyClass(IntWritable.class); job.setMapOutputValueClass(Text.class); // 如果Reducer输出类型是Text和LongWritable,保留以下设置 job.setOutputKeyClass(Text.class); job.setOutputValueClass(LongWritable.class);优化异常日志记录:
使用MapReduce日志工具或计数器记录异常,便于排查问题:import org.apache.commons.logging.Log; import org.apache.commons.logging.LogFactory; public static class MapClass extends Mapper<LongWritable,Text,IntWritable,Text> { private static final Log LOG = LogFactory.getLog(MapClass.class); public void map(LongWritable key, Text value, Context context) { try{ String[] str = value.toString().split(","); // 先校验数组长度 if(str.length < 2) { LOG.warn("行格式错误,列数不足:" + value.toString()); context.getCounter("MapperErrors", "InsufficientColumns").increment(1); return; } int int_year = Integer.parseInt(str[1]); context.write(new IntWritable(int_year),new Text(str[0])); } catch(Exception e) { LOG.error("处理行失败: " + value.toString(), e); context.getCounter("MapperErrors", "ParseErrors").increment(1); } } }处理表头行:
若数据包含表头,可通过判断行内容或偏移量跳过第一行:if(key.get() == 0) { // 跳过表头行 return; }
内容的提问来源于stack exchange,提问作者Aakash Rohila
相关产品推荐
相关产品推荐

