Hadoop MapReduce求航空公司最大总成本报错求助
Hey there! Let's break down the issues you're facing and fix them step by step—starting with that ls: "output" : No such file or directory error and moving to code corrections.
Key Issues in Your Code
1. Incorrect Command Line Argument Order
In your MaxTotalCost main class, you’re using args[1] as the input path and args[2] as the output path—but command line arguments start at index 0. This means you’re trying to read from a non-existent path (since args[2] doesn’t exist if you only pass two arguments) and writing to an invalid location, which is almost certainly causing the "output not found" error when you try to check results.
2. Mismatch Between Decimal Type and DoubleWritable
You mentioned the total cost is a Decimal type, but your code uses DoubleWritable which can introduce precision loss. For decimal values, we should use DecimalWritable (from org.apache.hadoop.io) to handle BigDecimal properly and preserve accuracy.
3. Unused Code in Reducer
The totalcostList in your reducer is never referenced—this is just unnecessary clutter that we can remove to clean up the code.
4. Fragile CSV Splitting
Splitting the CSV line with split(",") will break if any fields contain commas inside quotes (super common in CSV datasets). We’ll adjust the split logic to handle basic quoted fields for now.
Corrected Code
Mapper Class (Updated for Decimal Precision)
package org.myorg; import java.io.IOException; import java.math.BigDecimal; import org.apache.hadoop.io.DecimalWritable; import org.apache.hadoop.io.LongWritable; import org.apache.hadoop.io.Text; import org.apache.hadoop.mapreduce.Mapper; public class MaxTotalCostMapper extends Mapper<LongWritable, Text, Text, DecimalWritable> { private final static DecimalWritable totalcostWritable = new DecimalWritable(BigDecimal.ZERO); private Text airCarrier = new Text(); @Override public void map(LongWritable key, Text value, Context context) throws IOException, InterruptedException { // Skip the header line (adjust if your dataset doesn't have a header) if (key.get() == 0) { return; } // Split CSV line without breaking quoted fields String[] line = value.toString().split(",(?=(?:[^\"]*\"[^\"]*\")*[^\"]*$)"); // Ensure we have enough columns to avoid index out of bounds errors if (line.length >= 9) { airCarrier.set(line[8].trim()); try { BigDecimal totalCost = new BigDecimal(line[2].trim()); totalcostWritable.set(totalCost); context.write(airCarrier, totalcostWritable); } catch (NumberFormatException e) { // Track invalid cost values instead of crashing context.getCounter("MapperErrors", "InvalidTotalCost").increment(1); } } else { // Track malformed lines context.getCounter("MapperErrors", "MalformedLine").increment(1); } } }
Reducer Class (Cleaned Up)
package org.myorg; import java.io.IOException; import java.math.BigDecimal; import org.apache.hadoop.io.DecimalWritable; import org.apache.hadoop.io.Text; import org.apache.hadoop.mapreduce.Reducer; public class MaxTotalCostReducer extends Reducer<Text, DecimalWritable, Text, DecimalWritable> { @Override public void reduce(Text key, Iterable<DecimalWritable> values, Context context) throws IOException, InterruptedException { BigDecimal maxValue = BigDecimal.ZERO; for (DecimalWritable value : values) { BigDecimal current = value.get(); if (current.compareTo(maxValue) > 0) { maxValue = current; } } context.write(key, new DecimalWritable(maxValue)); } }
Main Class (Fixed Argument Order)
package org.myorg; import org.apache.hadoop.conf.Configuration; import org.apache.hadoop.fs.Path; import org.apache.hadoop.io.DecimalWritable; import org.apache.hadoop.io.Text; import org.apache.hadoop.mapreduce.Job; import org.apache.hadoop.mapreduce.lib.input.FileInputFormat; import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat; public class MaxTotalCost { public static void main(String[] args) throws Exception { Configuration conf = new Configuration(); if (args.length != 2) { System.err.println("Usage: MaxTotalCost <input path> <output path>"); System.exit(-1); } Job job = Job.getInstance(conf, "Max Total Cost"); job.setJarByClass(MaxTotalCost.class); // Fixed: args[0] = input path, args[1] = output path FileInputFormat.addInputPath(job, new Path(args[0])); FileOutputFormat.setOutputPath(job, new Path(args[1])); job.setMapperClass(MaxTotalCostMapper.class); job.setReducerClass(MaxTotalCostReducer.class); job.setOutputKeyClass(Text.class); job.setOutputValueClass(DecimalWritable.class); System.exit(job.waitForCompletion(true) ? 0 : 1); } }
Steps to Run Correctly
- Package the JAR: Export the corrected code as a JAR file from Eclipse.
- Run the Job: Use this Hadoop command (replace placeholders with your actual paths/JAR name):
Important: The output path must not exist before running—Hadoop will throw an error if it does.hadoop jar MaxTotalCost.jar org.myorg.MaxTotalCost /path/to/input/dft-flights-data-2011.csv /path/to/output - Check Results: After the job finishes, view the output with:
hdfs dfs -cat /path/to/output/part-r-00000
The ls: "output" : No such file or directory error was likely caused by the incorrect argument order in your main class (leading to the output path never being created) or the job failing silently due to other code issues. The corrected code fixes these problems and ensures your job runs as intended.
内容的提问来源于stack exchange,提问作者M Ravi

