如何通过YARN REST API部署Spark2作业并完成适配配置
Great question! Let's break this down step by step since Spark 2's YARN integration changed quite a bit from earlier versions—especially ditching the old spark-assembly.jar in favor of archived dependencies, which is exactly what your HDP cluster's spark2-hdp-yarn-archive.tar.gz is for.
1. Configuring am-container-spec Local Resources for Spark 2 Compatibility
Since Spark 2 no longer relies on a single monolithic assembly jar, you'll need to reference the pre-deployed spark2-hdp-yarn-archive.tar.gz in your local resources. This archive contains all the Spark 2 core dependencies, and YARN will automatically extract it into your container's working directory.
Here's how to structure this in your YARN REST API request's am-container-spec:
"am-container-spec": { "local-resources": [ { "resource": { "scheme": "hdfs", "host": "<your-nn-host>", "port": 8020, "file": "/path/to/spark2-hdp-yarn-archive.tar.gz" }, "type": "ARCHIVE", "visibility": "APPLICATION", "size": <file-size-in-bytes>, "timestamp": <file-timestamp>, "key": "spark2-hdp-yarn-archive" }, { // Add your application JAR here as a FILE resource "resource": { "scheme": "hdfs", "host": "<your-nn-host>", "port": 8020, "file": "/path/to/your-app.jar" }, "type": "FILE", "visibility": "APPLICATION", "size": <app-jar-size>, "timestamp": <app-jar-timestamp>, "key": "app.jar" } ], // ... rest of your container spec (commands, environment, etc.) }
Key notes:
- Set
typetoARCHIVEfor the Spark archive—this tells YARN to unpack it into a directory named after thekey(in this case,spark2-hdp-yarn-archive). visibility: "APPLICATION"ensures the resource is available to all containers in your application.
2. Building a REST-Compatible Classpath (Including Handling __spark.jar__)
In Spark 2, __spark.jar__ is no longer a single file—it's replaced by the collection of jars inside the extracted archive. You'll need to construct your classpath to include these jars, plus your application JAR, in the commands section of your container spec.
Here's how to set this up:
"am-container-spec": { // ... local-resources from above "commands": { "command": "\n export CLASSPATH=$PWD/app.jar:$PWD/spark2-hdp-yarn-archive/jars/*\n java -cp $CLASSPATH com.yourcompany.YourSparkMainClass\n " }, "environment": { "variables": { "SPARK_HOME": "$PWD/spark2-hdp-yarn-archive", "SPARK_YARN_MODE": "true" } } }
Important details:
- The
$PWD/spark2-hdp-yarn-archive/jars/*wildcard includes all Spark 2 core dependencies in the classpath. - We set
SPARK_HOMEto the extracted archive directory to align with Spark's expected environment variables. - You don't need to explicitly reference
__spark.jar__anymore—this placeholder was specific to the old assembly jar approach, which Spark 2 abandons.
3. Reducing Duplicate Configuration by Mirroring Java YARN API Patterns
The Java YARN API abstracts away a lot of repetitive setup, and you can replicate this pattern for your REST API submissions:
- Create reusable configuration snippets: Define a base JSON template for the
am-container-specthat includes the Spark archive resource, environment variables, and classpath structure. Each time you submit a job, just swap out the application JAR, main class, and any job-specific parameters. - Parameterize common values: Use variables for things like HDFS namenode host/port, archive path, default resource sizes (memory/CPU), so you don't have to hardcode them every time.
- Script the request generation: Write a simple shell or Python script that takes job-specific inputs (app jar path, main class, args) and injects them into your base template. This mirrors how the Java API's
SparkYarnClientbuilds the YARN application context programmatically.
For example, a shell script could use jq to patch a base template with job-specific values:
# Base template path BASE_TEMPLATE="./spark2-yarn-rest-base.json" # Job-specific values APP_JAR="/user/rick/my-app.jar" MAIN_CLASS="com.rick.MySparkJob" # Generate final request JSON jq --arg jar "$APP_JAR" \ --arg main "$MAIN_CLASS" \ '.am-container-spec.local-resources[1].resource.file = $jar | .am-container-spec.commands.command = "\n export CLASSPATH=$PWD/app.jar:$PWD/spark2-hdp-yarn-archive/jars/*\n java -cp $CLASSPATH \($main)\n "' \ $BASE_TEMPLATE > final-request.json
This approach cuts down on repetitive configuration and reduces the chance of typos or missing settings.
内容的提问来源于stack exchange,提问作者Rick Moritz

