关于YARN框架下MapReduce作业阶段的技术咨询
Hey there! Awesome that you’ve got a solid grasp of YARN fundamentals—let’s break down MapReduce’s key stages and how it fits into YARN’s multi-workload ecosystem clearly.
YARN & MapReduce: The Big Picture
First, a quick refresher to tie things together: YARN acts as the cluster’s resource manager, handling CPU, memory, and storage allocation across all nodes. MapReduce is one of the many compute frameworks that runs on top of YARN (alongside Spark, Flink, etc.).
Each framework gets its own ApplicationMaster when a job is submitted—for MapReduce, this is the MRAppMaster. The MRAppMaster works with YARN’s ResourceManager to request containers (the unit of resource allocation) for Mapper and Reducer tasks, while NodeManagers on each cluster node handle actually running those tasks. This is why YARN can run multiple types of jobs side by side: it’s agnostic to the compute logic, just managing resources.
MapReduce Core Stages Breakdown
Let’s walk through each critical phase of a MapReduce job, step by step:
1. Mapper Phase
- What it does: Takes raw input data (split into chunks called splits, usually aligned with HDFS blocks) and transforms it into key-value pairs (
K2, V2) via your custommap()function. - Parallelism: Each input split gets its own Mapper task—YARN spins up containers across cluster nodes to run these in parallel, maximizing throughput.
- Example: If you’re processing web logs, a Mapper might take a log line like
192.168.1.1 - [10/Oct/2000:13:55:36 -0700] "GET /index.html HTTP/1.0"and output(192.168.1.1, 1)to count visits per IP. - Under the hood: Mapper outputs are first written to local disk (not HDFS) to avoid network overhead, then prepped for the next phase.
2. Sorting & Shuffle Phase
This is often the most confusing part—let’s unpack it:
- Shuffle: The process of transferring Mapper outputs to the correct Reducer tasks. Mapper outputs are partitioned by key (using a default hash partitioner, or a custom one you define), so all values for the same key go to the same Reducer.
- Sorting: Happens in two places:
- Mapper-side: After writing outputs to local disk, MapReduce automatically sorts them by key.
- Reducer-side: As Reducers pull partitioned data from Mappers, they merge and sort the incoming key-value pairs to group all values for a single key together.
- Why this matters: Sorting ensures Reducers receive all related data in a structured way, making aggregation straightforward.
3. Reducer Phase
- What it does: Takes the sorted, grouped key-value pairs from the Shuffle phase and runs your custom
reduce()function to aggregate, filter, or transform the data into final results (K3, V3). - Parallelism: The number of Reducers is configurable (you set this when submitting the job). YARN allocates containers for each Reducer, which run in parallel across the cluster.
- Example: Using the web log example, a Reducer would receive
(192.168.1.1, [1, 1, 1])and output(192.168.1.1, 3)to show the total visits for that IP. - Under the hood: Final Reducer outputs are written directly to HDFS (or your configured output storage) for persistence.
Quick Workflow Recap with YARN
To tie it all to YARN’s role:
- You submit a MapReduce job to YARN’s ResourceManager.
- ResourceManager allocates a container to launch the
MRAppMaster. - MRAppMaster calculates how many Mapper/Reducer tasks are needed, then requests containers from ResourceManager.
- NodeManagers start the Mapper/Reducer containers and execute the tasks.
- Once all tasks finish, MRAppMaster reports job completion to ResourceManager, which releases all allocated resources.
内容的提问来源于stack exchange,提问作者CuriousMind

