启动Jupyter Notebook服务器时如何设置PYTHONPATH以使用PySpark
Fixing PySpark 2.2.1 Import in Jupyter Public Server
Got it, let's tackle this issue—your local Python can access PySpark because it loads the environment variables from .bashrc, but your Jupyter public server doesn't pick those up when starting. Here are three solid solutions to get PySpark working in your notebooks:
1. Add Environment Variables to Jupyter's Configuration File
This makes Jupyter load the Spark variables every time it starts.
- First, generate a Jupyter config file if you don't have one already:
jupyter notebook --generate-config - Open the config file (usually located at
~/.jupyter/jupyter_notebook_config.py) in your editor. - Add these lines at the top (replace the paths with your actual Spark installation directory):
import os # Set Spark home directory os.environ['SPARK_HOME'] = '/path/to/your/spark-2.2.1' # Configure Python path to include Spark's Python library and py4j spark_python_path = os.path.join(os.environ['SPARK_HOME'], 'python') py4j_path = os.path.join(os.environ['SPARK_HOME'], 'python/lib/py4j-0.10.4-src.zip') os.environ['PYTHONPATH'] = f"{spark_python_path}:{py4j_path}:{os.environ.get('PYTHONPATH', '')}" - Restart your Jupyter server, then test importing PySpark in a notebook.
2. Create a Custom PySpark Kernel
This lets you select a dedicated kernel for PySpark with preloaded environment variables.
- Generate a new kernel profile:
ipython kernel install --name pyspark22 --user - Find the kernel's JSON config file (typically at
~/.local/share/jupyter/kernels/pyspark22/kernel.json). - Replace its content with this (update paths to match your setup):
{ "argv": [ "python", "-m", "ipykernel_launcher", "-f", "{connection_file}" ], "display_name": "PySpark 2.2.1", "env": { "SPARK_HOME": "/path/to/your/spark-2.2.1", "PYTHONPATH": "/path/to/your/spark-2.2.1/python:/path/to/your/spark-2.2.1/python/lib/py4j-0.10.4-src.zip:$PYTHONPATH", "PYSPARK_PYTHON": "python" } } - Restart Jupyter, and you'll see the "PySpark 2.2.1" kernel in your notebook's kernel selector. Use this kernel for PySpark work.
3. Update Jupyter Server Startup Script (For Service-Managed Servers)
If you're running Jupyter as a system service (like systemd), set the environment variables directly in the service config:
- Open your Jupyter service file (e.g.,
/etc/systemd/system/jupyter.service). - Add these lines under the
[Service]section:Environment="SPARK_HOME=/path/to/your/spark-2.2.1" Environment="PYTHONPATH=/path/to/your/spark-2.2.1/python:/path/to/your/spark-2.2.1/python/lib/py4j-0.10.4-src.zip:$PYTHONPATH" - Reload the systemd daemon and restart the Jupyter service:
sudo systemctl daemon-reload sudo systemctl restart jupyter
Quick Verification
After applying any of these methods, run this in a notebook to confirm the variables are loaded:
import os print("SPARK_HOME:", os.environ.get("SPARK_HOME")) print("PYTHONPATH:", os.environ.get("PYTHONPATH"))
If the paths are correct, you should be able to import pyspark without errors.
内容的提问来源于stack exchange,提问作者Salva Martini

