Hadoop ToolRunner接口使用异常:仍出现GenericOptionsParser警告
我已经实现了ToolRunner接口,但执行Hadoop任务时仍收到WARN [JobClient] Use GenericOptionsParser警告。执行命令如下:
hadoop jar gc.jar stubs.AvgWordLength -D mapred.reduce.tasks=10 myInput.txt myOutput_res
命令可正常运行,但警告无法消除。使用Linux系统,Driver/ToolRunner代码如下:
package stubs; import org.apache.hadoop.fs.Path; import org.apache.hadoop.io.DoubleWritable; import org.apache.hadoop.io.IntWritable; import org.apache.hadoop.io.Text; import org.apache.hadoop.mapreduce.lib.input.FileInputFormat; import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat; import org.apache.hadoop.mapreduce.Job; import org.apache.hadoop.conf.Configured; import org.apache.hadoop.conf.Configuration; import org.apache.hadoop.util.Tool; import org.apache.hadoop.util.ToolRunner; public class AvgWordLength extends Configured implements Tool { public static void main(String[] args) throws Exception { int exitCode = ToolRunner.run(new Configuration(), new AvgWordLength(), args); System.exit(exitCode); } @Override public int run(String[] args) throws Exception { /* * Validate that two arguments were passed from the command line. */ if (args.length != 2) { System.out.printf("Usage: AvgWordLength <input dir> <output dir>\n"); System.exit(-1); } /* * Instantiate a Job object for your job's configuration. */ Job job = Job.getInstance(getConf()); /* * Specify the jar file that contains your driver, mapper, and reducer. * Hadoop will transfer this jar file to nodes in your cluster running * mapper and reducer tasks. */ job.setJarByClass(AvgWordLength.class); /* * Specify an easily-decipherable name for the job. This job name will * appear in reports and logs. */ job.setJobName("Average Word Length"); /* * Specify the paths to the input and output data based on the * command-line arguments. */ FileInputFormat.setInputPaths(job, new Path(args[0])); FileOutputFormat.setOutputPath(job, new Path(args[1])); /* * Specify the mapper and reducer classes. */ job.setMapperClass(LetterMapper.class); job.setReducerClass(AverageReducer.class); /* * The input file and output files are text files, so there is no need * to call the setInputFormatClass and setOutputFormatClass methods. */ /* * The mapper's output keys and values have different data types than * the reducer's output keys and values. Therefore, you must call the * setMapOutputKeyClass and setMapOutputValueClass methods. */ job.setMapOutputKeyClass(Text.class); job.setMapOutputValueClass(IntWritable.class); /* * Specify the job's output key and value classes. */ job.setOutputKeyClass(Text.class); job.setOutputValueClass(DoubleWritable.class); /* * Start the MapReduce job and wait for it to finish. If it finishes * successfully, return 0. If not, return 1. */ boolean success = job.waitForCompletion(true); return (success ? 0 : 1); } }
这个警告的核心原因是代码中Configuration的传递方式不符合Hadoop对通用参数解析的预期,尽管你使用了ToolRunner,但手动传入的new Configuration()干扰了参数解析流程,具体修改如下:
- 调整main方法中的ToolRunner调用
将main方法中的调用代码从:
int exitCode = ToolRunner.run(new Configuration(), new AvgWordLength(), args);
修改为:
int exitCode = ToolRunner.run(new AvgWordLength(), args);
因为AvgWordLength继承了Configured,ToolRunner会自动为其创建并初始化Configuration,同时自动解析命令行中的通用参数(如-D开头的配置项)到这个Configuration中,无需手动传入新的Configuration实例。
确认配置传递正确性
run方法中Job.getInstance(getConf())的写法是正确的,getConf()会返回ToolRunner处理过的、已经包含命令行参数的Configuration对象,确保你的reduce任务数等配置能正常生效。关于警告的额外说明
如果修改后仍出现该警告,大概率是Hadoop旧版本的兼容性问题——部分早期Hadoop版本即使正确使用ToolRunner,也会弹出这个提示。此时可以忽略警告,只要你的任务能正常运行、配置参数生效(可通过Hadoop Web UI查看任务的reduce数量是否为10)即可。
内容的提问来源于stack exchange,提问作者Bernard Joshua

