Map输出键类型不匹配问题:预期org.apache.hadoop.io.Text,实际收到org.apache.hadoop.io.LongWritable,如何解决?
Hey there! Let's work through this type mismatch issue and get your Hadoop job up and running. I'll break down the problems in your code and show you how to fix them, plus add the missing Reduce step you need to calculate total cases per continent.
First, Let's Address the Type Mismatch Errors
Your code has a few key issues that are causing the type mismatch:
Incorrect Mapper Generic Type Order
Hadoop's defaultTextInputFormatpasses input as:- Key:
LongWritable(the byte offset of the current line) - Value:
Text(the full line of text from the CSV)
You flipped these in your Mapper declaration—let's fix that.
- Key:
Wrong Method Name & API Usage
- The Mapper's overridden method is lowercase
map(), not uppercaseMap()(Java is case-sensitive!). - You're using the old
mapredAPI'sOutputCollector, but your code imports the newermapreduceAPI. We need to useContextinstead to emit key-value pairs.
- The Mapper's overridden method is lowercase
Missing Reduce Class
You mentioned you want to calculate cumulative totals, but you haven't implemented a Reducer yet. The Map step emits (continent, new cases) pairs, and the Reduce step will sum those values per continent.
Modified Map Class Code
package ContTotCase; import java.io.IOException; import org.apache.hadoop.io.LongWritable; import org.apache.hadoop.io.Text; import org.apache.hadoop.mapreduce.Mapper; public class ContTotCasesMap extends Mapper<LongWritable, Text, Text, LongWritable> { @Override protected void map(LongWritable key, Text value, Context context) throws IOException, InterruptedException { // Split the CSV line (note: this doesn't handle commas inside quotes, but works for simple CSVs) String[] fields = value.toString().split(","); // Make sure we have at least 6 fields (indexes 0-5) to avoid ArrayIndexOutOfBoundsException if (fields.length >= 6) { String continent = fields[1].trim(); String caseStr = fields[5].trim(); // Try to parse the case count, skip invalid entries try { long cases = Long.parseLong(caseStr); if (cases > 0) { context.write(new Text(continent), new LongWritable(cases)); } } catch (NumberFormatException e) { // Skip lines where the case count isn't a valid number System.err.println("Skipping invalid case count: " + caseStr); } } } }
Add a Reduce Class
Create this new class to sum the cases per continent:
package ContTotCase; import java.io.IOException; import org.apache.hadoop.io.LongWritable; import org.apache.hadoop.io.Text; import org.apache.hadoop.mapreduce.Reducer; public class ContTotCasesReduce extends Reducer<Text, LongWritable, Text, LongWritable> { @Override protected void reduce(Text key, Iterable<LongWritable> values, Context context) throws IOException, InterruptedException { long total = 0; // Sum all the case values for this continent for (LongWritable val : values) { total += val.get(); } // Emit the continent and its total cases context.write(key, new LongWritable(total)); } }
Updated Driver Class
Now update your Driver to include the Reducer and fix any minor issues:
package ContTotCase; import org.apache.hadoop.conf.Configuration; import org.apache.hadoop.fs.Path; import org.apache.hadoop.io.LongWritable; import org.apache.hadoop.io.Text; import org.apache.hadoop.mapreduce.Job; import org.apache.hadoop.mapreduce.lib.input.FileInputFormat; import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat; public class ContTotCaseDriver { public static void main(String[] args) throws Exception { if (args.length != 2) { System.err.println("Usage: ContTotCases <Input Path> <Output Path>"); System.exit(2); } Configuration conf = new Configuration(); Job job = Job.getInstance(conf, "Continent Total COVID Cases"); // Set all required classes job.setJarByClass(ContTotCaseDriver.class); job.setMapperClass(ContTotCasesMap.class); job.setReducerClass(ContTotCasesReduce.class); // Add this line to use our Reducer // Set output types (map output can match final output here) job.setOutputKeyClass(Text.class); job.setOutputValueClass(LongWritable.class); FileInputFormat.addInputPath(job, new Path(args[0])); FileOutputFormat.setOutputPath(job, new Path(args[1])); System.exit(job.waitForCompletion(true) ? 0 : 1); } }
Key Notes for You as a Java/Hadoop Newbie
- Always use
@Overridewhen overriding methods (likemap()andreduce())—it helps catch typos (like your uppercaseMap()mistake). - Handle edge cases like missing CSV fields or invalid numbers to avoid crashes.
- The Reducer runs once per unique key (continent), and gets all the values emitted for that key from the Mapper. Summing those gives your total.
内容的提问来源于stack exchange,提问作者Leuz

