Spark with YARN集群能否运行Python/R脚本及最简运行方式咨询
Hey there! Great question—let's break this down clearly:
Support for Python/R Scripts
Spark with YARN fully supports Python (via PySpark) and R (via SparkR or Sparklyr) scripts for big data workloads. A quick caveat though: it's not every arbitrary local Python/R script that works out of the box.
- If your script uses Spark's distributed computing APIs (the core reason for using Spark in big data), it will run seamlessly on YARN.
- Even if you have a plain Python/R script with no Spark logic, you can still leverage YARN's resource scheduling to run it across the cluster—you just need to package/submit it correctly.
Simplest Run Methods
Assuming you have a Spark-compatible script or want to run a basic script on YARN, here are the minimal steps for each language:
1. Python (PySpark or Plain Script)
Example PySpark Script (simple_pyspark.py)
from pyspark.sql import SparkSession if __name__ == "__main__": # Initialize Spark session with YARN as master spark = SparkSession.builder.appName("PySpark-YARN-Demo").getOrCreate() # Sample workload: Read a text file and display content df = spark.read.text("/user/your-name/sample-input.txt") df.show() spark.stop()
Minimal Submit Command
Run this on any cluster node with the Spark client installed:
spark-submit --master yarn --deploy-mode client simple_pyspark.py
--master yarn: Tells Spark to use YARN for resource management--deploy-mode client: Runs the driver program on your local machine (great for debugging; useclustermode for production to run the driver on a cluster node)
For a plain Python script with no Spark dependencies, you can still use the same command—just adjust the executor count as needed:
spark-submit --master yarn --deploy-mode client --num-executors 1 your-plain-python-script.py
2. R (SparkR or Sparklyr)
Example SparkR Script (simple_sparkr.R)
# Load SparkR library library(SparkR) # Initialize Spark session connected to YARN sparkR.session(appName = "SparkR-YARN-Demo", master = "yarn") # Sample workload: Read and display a text file df <- read.text("/user/your-name/sample-input.txt") showDF(df) # Clean up session sparkR.session.stop()
Minimal Submit Command
Same pattern as Python—run this on a Spark client node:
spark-submit --master yarn --deploy-mode client simple_sparkr.R
If you use Sparklyr instead of SparkR, just ensure your script specifies master = "yarn" when initializing the connection, then submit with the same spark-submit command.
Quick Notes
- Make sure your cluster has Python/R installed (matching the versions your Spark is configured for)
- Set
spark.pyspark.python(for Python) orspark.r.command(for R) in your Spark config if the default paths aren't correct - For third-party dependencies, use
--py-files(Python zip/egg files) or--packages(Maven dependencies) withspark-submit, or pre-install packages on all cluster nodes
内容的提问来源于stack exchange,提问作者DJ R

