Hadoop MapReduce Reducer拆分Text Key触发数组越界异常问题
问题原因分析与解决方案
异常根源
你遇到的java.lang.ArrayIndexOutOfBoundsException: 1,核心原因是Reducer接收到的部分key不符合「xxx-xxx」的格式——当调用split("-")时,这些key无法分割出两个元素,导致数组长度为1,此时访问索引1自然就触发了越界异常。
具体看你的Mapper代码:在catch (NumberFormatException e)代码块中,你写入的key是new Text("perso"),这个字符串里没有-分隔符。当这条数据传到Reducer后,key.toString().split("-")得到的数组只有["perso"]一个元素,长度为1,这时你直接访问out[1]必然抛出异常。
修复方案
你可以从两个方向解决这个问题:
方案1:在Reducer中做格式校验(推荐)
在访问split后的数组元素前,先判断数组长度是否满足要求,避免越界:
public void reduce(Text key, Iterable<MapWritable> values, Context context) throws IOException, InterruptedException { out = key.toString().split("-"); // 先校验数组长度,确保有两个元素 if (out.length >= 2) { for (MapWritable m : values) { context.write(new Text(out[1]), m); } } else { // 可选:统计这类异常key的数量,方便排查问题 context.getCounter("Reducer_Error", "Invalid_Key_Format").increment(1); // 或者直接跳过该key,不进行输出 return; } }
方案2:统一Mapper输出的key格式
在Mapper的catch块中,也生成符合「xxx-xxx」格式的key,比如:
catch (NumberFormatException e) { // 给异常数据也加上分隔符,保证格式统一 context.write(new Text("perso-0"), missing); }
这样所有Reducer接收到的key都能split出至少两个元素,不会触发越界。
额外小建议
你的Mapper中mw、missing是类成员变量,而MapReduce的Mapper实例会被多线程复用,这可能导致不同map任务的数据互相污染(比如前一次的mw数据被后续任务复用)。建议把这些变量放到map方法内部声明,每次调用都创建新实例:
public void map(LongWritable key, Text value, Context context) throws IOException, InterruptedException { // 把变量移到方法内,避免线程安全问题 MapWritable missing = new MapWritable(); missing.put(new LongWritable(0), new IntWritable(0)); MapWritable mw = new MapWritable(); ArrayWritable aw = new ArrayWritable(IntWritable.class); int[] score; // ... 后续业务代码 }
内容的提问来源于stack exchange,提问作者Gianlorenzo Didonato
相关产品推荐
相关产品推荐

