运行Hadoop MapReduce WordCount报java.io.IOException: Job failed原因排查
Hadoop MapReduce Job失败原因排查
报错异常:
java.io.IOException: Job failed! at org.apache.hadoop.mapred.JobClient.runJob(JobClient.java:873) at WordCount1Driver.main(WordCount1Driver.java:45)
1. Mapper数组下标越界
你在WordCount1Mapper类的map方法中直接调用SingleCountryData[7]获取按空格拆分后的第8个元素:
String[] SingleCountryData = valueString.split(" "); output.collect(new Text(SingleCountryData[7]), one);
如果输入文件中某一行的空格分隔后的元素总数小于8,会直接抛出ArrayIndexOutOfBoundsException,导致Map任务终止,最终Job运行失败。
修复方案:取元素前先判断数组长度:
String[] SingleCountryData = valueString.split(" "); if (SingleCountryData.length >= 8) { output.collect(new Text(SingleCountryData[7]), one); }
2. Java路径转义错误
你在WordCount1Driver中硬编码的Windows本地路径使用了单反斜杠:
FileInputFormat.setInputPaths(job_conf, new Path("D:\program files\eclipse\WordCount\input.txt")); FileOutputFormat.setOutputPath(job_conf, new Path("D:\program files\eclipse\WordCount\output"));
Java中反斜杠\是转义字符,单写会被识别为特殊转义符,导致路径解析错误,无法找到输入文件或者无法创建输出目录。
修复方案:路径改用双反斜杠或者正斜杠:
// 方式1:双反斜杠转义 FileInputFormat.setInputPaths(job_conf, new Path("D:\\program files\\eclipse\\WordCount\\input.txt")); FileOutputFormat.setOutputPath(job_conf, new Path("D:\\program files\\eclipse\\WordCount\\output")); // 方式2:使用正斜杠(Java兼容Windows和Linux的写法) FileInputFormat.setInputPaths(job_conf, new Path("D:/program files/eclipse/WordCount/input.txt")); FileOutputFormat.setOutputPath(job_conf, new Path("D:/program files/eclipse/WordCount/output"));
3. 输出目录已存在
Hadoop MapReduce的设计要求输出目录必须是运行前不存在的目录,如果你之前已经运行过程序生成了output目录,再次运行会直接报错导致任务失败。
修复方案:每次运行前手动删除output目录,或者在Driver代码中加入自动判断删除逻辑:
// 加入到JobClient.runJob(job_conf);之前 Path outputPath = new Path("D:/program files/eclipse/WordCount/output"); FileSystem fs = FileSystem.get(job_conf); if(fs.exists(outputPath)){ fs.delete(outputPath, true); }
其他非致命逻辑问题
你当前的Mapper逻辑存在重复输出的问题:先输出拆分后的第8个字段,又遍历输出整行的所有分词,会导致最终统计结果不符合预期,可根据你的实际业务需求保留其中一套逻辑即可。
内容的提问来源于stack exchange,提问作者Lee Huy
相关产品推荐
相关产品推荐

