YARN如何确定启动对应Application Master?RM如何识别提交的作业类型?
Great question! Let's break down the concrete, step-by-step mechanisms that let YARN's ResourceManager (RM) know exactly which ApplicationMaster (AM) to spin up when you submit jobs like MapReduce or Spark.
1. The Client Does the "Type Identification" Upfront
The RM doesn't have to guess what kind of job you're submitting—your client explicitly tells it which AM to run. When you submit a job:
- For MapReduce: The
hadoop jarcommand (or your MapReduce client code) sets the configuration propertyyarn.appmaster.main.classtoorg.apache.hadoop.mapreduce.v2.app.MRAppMaster - For Spark: The
spark-submittool (with--master yarn) sets the same property toorg.apache.spark.deploy.yarn.ApplicationMaster
This fully qualified class name is the direct "identifier" the RM relies on—it's hard-coded into the client's submission logic for each framework.
2. Critical Details Are Passed via ApplicationSubmissionContext
When submitting a job, the client communicates with the RM using YARN's ApplicationClientProtocol API. It sends an ApplicationSubmissionContext object that includes:
- The exact AM main class name we mentioned above
- The shell command to launch the AM process on a NodeManager
- Resource requirements (memory, vcores) for the AM's container
- Classpath settings to ensure the AM's dependencies are available on the target node
The RM's ApplicationTrackerService parses this context immediately—this is how it gets all the info it needs to start the right AM.
3. RM Launches the AM Without Needing to Understand Job Logic
The RM only handles resource allocation—it doesn't care what the AM actually does once it's running. Here's the simplified flow:
- RM allocates the first container (reserved exclusively for the AM) based on the resource specs in the submission context
- It sends a command to the target NodeManager to launch a process using the specified AM main class
- Once the AM starts, it takes over: it negotiates with the RM for additional containers (Map/Reduce tasks for MR, Executors for Spark) and manages the full job execution lifecycle
Quick Example: Spark Job Submission Under the Hood
When you run spark-submit --master yarn --class com.yourteam.YourSparkApp your-app.jar:
spark-submitbuilds theApplicationSubmissionContextwithorg.apache.spark.deploy.yarn.ApplicationMasteras the AM class- It sends this context to the RM
- RM allocates an AM container and tells the NodeManager to start the Spark AM
- The Spark AM then requests Executor containers from the RM and runs your application
内容的提问来源于stack exchange,提问作者CuriousMind

