You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

已实现Jupyter本地运行PySpark,如何切换至YARN-client等多模式?

Switching PySpark Execution Modes in Jupyter

Great to hear you've got PySpark running locally in Jupyter! Let's walk through how to configure it for YARN-client, standalone, and YARN-cluster modes, with practical code snippets and key notes for each.

1. YARN-client Mode (Interactive-Friendly)

YARN-client mode keeps the PySpark driver running directly in your Jupyter process, which makes it perfect for interactive work—you'll see logs and results right in your notebook, and can debug in real time.

Prerequisites

  • Your Hadoop YARN cluster is up and running, and your Jupyter machine has network access to the YARN ResourceManager.
  • Ensure Hadoop configuration files (core-site.xml, yarn-site.xml, etc.) are either in Spark's conf directory, or you've set the HADOOP_CONF_DIR environment variable to point to their location.

Code Configuration

import findspark
# Uncomment and set the path if SPARK_HOME isn't in your environment variables
# findspark.init("/path/to/your/spark/install")
from pyspark import SparkContext, SparkConf

# Build the configuration for YARN-client mode
conf = SparkConf() \
    .setAppName("My-YARN-Client-App") \
    .setMaster("yarn") \
    .set("spark.submit.deployMode", "client")

# Initialize the SparkContext
sc = SparkContext(conf=conf)

2. Standalone Mode (Spark's Own Cluster)

Standalone mode uses Spark's built-in cluster manager. First, you'll need to start the Spark standalone cluster (run start-all.sh from your Spark sbin directory on the master node).

Code Configuration

import findspark
findspark.init()
from pyspark import SparkContext, SparkConf

# Replace <master-node-ip> with your Spark Master's IP address (default port is 7077)
conf = SparkConf() \
    .setAppName("My-Standalone-App") \
    .setMaster("spark://<master-node-ip>:7077")

# Initialize the SparkContext (default deploy mode is client, which works for Jupyter)
sc = SparkContext(conf=conf)

Note

Standalone cluster mode (where the driver runs on a worker node) isn't ideal for Jupyter, since you lose direct interactive access to the driver. Stick with client mode for notebook-based work.

3. YARN-cluster Mode (Production-Focused)

YARN-cluster mode runs the PySpark driver on a YARN NodeManager node in the cluster. This is great for production jobs, but not recommended for interactive Jupyter sessions—since the driver isn't local, your notebook can't directly communicate with it to return results or accept commands.

Code (For Reference)

If you still want to test it (though expect limited interactivity):

import findspark
findspark.init()
from pyspark import SparkContext, SparkConf

conf = SparkConf() \
    .setAppName("My-YARN-Cluster-App") \
    .setMaster("yarn") \
    .set("spark.submit.deployMode", "cluster")

sc = SparkContext(conf=conf)

Better Alternative for YARN-cluster

For production jobs, use spark-submit instead of Jupyter:

spark-submit --master yarn --deploy-mode cluster your_pyspark_script.py

内容的提问来源于stack exchange,提问作者Shengxin Huang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:18:02