已实现Jupyter本地运行PySpark,如何切换至YARN-client等多模式?
Great to hear you've got PySpark running locally in Jupyter! Let's walk through how to configure it for YARN-client, standalone, and YARN-cluster modes, with practical code snippets and key notes for each.
1. YARN-client Mode (Interactive-Friendly)
YARN-client mode keeps the PySpark driver running directly in your Jupyter process, which makes it perfect for interactive work—you'll see logs and results right in your notebook, and can debug in real time.
Prerequisites
- Your Hadoop YARN cluster is up and running, and your Jupyter machine has network access to the YARN ResourceManager.
- Ensure Hadoop configuration files (
core-site.xml,yarn-site.xml, etc.) are either in Spark'sconfdirectory, or you've set theHADOOP_CONF_DIRenvironment variable to point to their location.
Code Configuration
import findspark # Uncomment and set the path if SPARK_HOME isn't in your environment variables # findspark.init("/path/to/your/spark/install") from pyspark import SparkContext, SparkConf # Build the configuration for YARN-client mode conf = SparkConf() \ .setAppName("My-YARN-Client-App") \ .setMaster("yarn") \ .set("spark.submit.deployMode", "client") # Initialize the SparkContext sc = SparkContext(conf=conf)
2. Standalone Mode (Spark's Own Cluster)
Standalone mode uses Spark's built-in cluster manager. First, you'll need to start the Spark standalone cluster (run start-all.sh from your Spark sbin directory on the master node).
Code Configuration
import findspark findspark.init() from pyspark import SparkContext, SparkConf # Replace <master-node-ip> with your Spark Master's IP address (default port is 7077) conf = SparkConf() \ .setAppName("My-Standalone-App") \ .setMaster("spark://<master-node-ip>:7077") # Initialize the SparkContext (default deploy mode is client, which works for Jupyter) sc = SparkContext(conf=conf)
Note
Standalone cluster mode (where the driver runs on a worker node) isn't ideal for Jupyter, since you lose direct interactive access to the driver. Stick with client mode for notebook-based work.
3. YARN-cluster Mode (Production-Focused)
YARN-cluster mode runs the PySpark driver on a YARN NodeManager node in the cluster. This is great for production jobs, but not recommended for interactive Jupyter sessions—since the driver isn't local, your notebook can't directly communicate with it to return results or accept commands.
Code (For Reference)
If you still want to test it (though expect limited interactivity):
import findspark findspark.init() from pyspark import SparkContext, SparkConf conf = SparkConf() \ .setAppName("My-YARN-Cluster-App") \ .setMaster("yarn") \ .set("spark.submit.deployMode", "cluster") sc = SparkContext(conf=conf)
Better Alternative for YARN-cluster
For production jobs, use spark-submit instead of Jupyter:
spark-submit --master yarn --deploy-mode cluster your_pyspark_script.py
内容的提问来源于stack exchange,提问作者Shengxin Huang

