Hadoop MapReduce中cleanup()方法工作原理及Reducer代码疑问
Hey there! Let's unpack your questions one by one—this is a classic top-N pattern in MapReduce, so I totally get why this might feel confusing at first.
First: What's the if (Top5DataEngineer.size() > 1) ... line doing?
Let's start with context: Your Top5DataEngineer is a TreeMap<LongWritable, Text>, which sorts entries in ascending order of the key (since LongWritable follows natural numeric order).
Here's the play-by-play for each reduce call:
- For each input key (which is a combination of
Region\tYear, like"XYZ\t2016"), you calculate the total job countsumfor that region in the given year. - You add this result to the TreeMap, using
sumas the key and the formatted string (key + "," + sum) as the value. - The
ifcheck: If the TreeMap has more than 1 entry, you remove the entry with the smallest key (viafirstKey()).
Since your Partitioner sends all data for a single year to one Reducer, this logic ensures the TreeMap always keeps only the largest job count entry encountered so far for that year. By the time all reduce calls finish, the TreeMap holds exactly the region with the highest job count for the year.
(Side note: The variable name Top5DataEngineer is a bit misleading here—your code is actually keeping only the top 1 entry per year, not top 5. If you wanted top 5, you'd change the condition to >5 instead of >1.)
Second: How does the cleanup() method work?
MapReduce Reducers have a defined lifecycle per Reducer instance:
- The Reducer is initialized (via
setup(), which you aren't using here). - The
reduce()method is called once for each unique key in the Reducer's input partition. - After all
reduce()calls are completed, thecleanup()method runs exactly once per Reducer instance.
Its purpose is for final cleanup or output tasks—like writing the final aggregated results that you've been collecting during all the reduce() calls.
Why moving cleanup code to reduce() breaks things?
If you put the loop that writes the TreeMap entries into reduce(), here's what happens:
- Every time
reduce()processes a region's key, you'd output whatever is in the TreeMap at that moment. - For example, if you have 10 regions in a year, you'd output 10 times—each time showing the "current top" entry up to that point. This would result in duplicate or incorrect output, not just the final top region.
By keeping the output in cleanup(), you wait until all regions for the year have been processed, then output the single, final top entry that's left in the TreeMap—exactly what you want.
Quick Recap of Your Code's Flow
To tie it all together:
- Mapper filters for Data Engineer roles, outputs keys like
Region\tYearand value1. - Partitioner sends all data for a single year to one Reducer (2011→Reducer 0, 2012→Reducer1, etc.).
- Each Reducer uses the TreeMap to track the highest job count region as it processes each region's data.
- After all regions are processed,
cleanup()writes that single top entry for the year.
That's why your code works when using cleanup() but not when moving the output to reduce()!
内容的提问来源于stack exchange,提问作者Anand Raina

