Jupyter Notebook与PySpark集成调试求助:本地环境配置后运行报错
Hey there! Let’s work through this frustrating PySpark + Jupyter problem you’re hitting. That TypeError: an integer is required (got type bytes) error is a classic sign of a Python version mismatch with your Spark installation, and once we fix that, we’ll get your local environment up and running for AWS Glue testing in no time.
Root Cause
You’re using Spark 2.4.8, which was released before Python 3.8 came out. Python 3.8 changed how the types.CodeType constructor works (it now expects integers for certain parameters instead of bytes), and the older cloudpickle library bundled with Spark 2.4.x doesn’t handle this change. That’s exactly what’s causing the stack trace you shared.
Step-by-Step Solutions
1. Create a Python 3.7 Conda Environment (Critical!)
Spark 2.4.x officially supports Python 3.6 and 3.7—skip 3.8+ entirely for this setup. Here’s how to set up a compatible environment:
- Open Anaconda Prompt and run:
conda create -n pyspark_glue_env python=3.7 conda activate pyspark_glue_env - Install required packages for your workflow:
conda install jupyter findspark pip install aws-glue-libs # For AWS Glue code testing
2. Verify & Update Environment Variables
Double-check these system environment variables (adjust paths to match your installations):
SPARK_HOME: Points to your Spark 2.4.8 directory (e.g.,C:\spark-2.4.8-bin-hadoop2.7)HADOOP_HOME: Points to your winutils directory (e.g.,C:\winutils\hadoop-2.7.1)- Add
%SPARK_HOME%\binand%HADOOP_HOME%\binto your System PATH
For your new Conda environment, set these session-specific variables to ensure PySpark uses the correct Python interpreter:
set PYSPARK_PYTHON=C:\Users\[YourUsername]\anaconda3\envs\pyspark_glue_env\python.exe set PYSPARK_DRIVER_PYTHON=jupyter set PYSPARK_DRIVER_PYTHON_OPTS=notebook
3. Test the Fixed Environment
With your Conda environment activated and variables set:
- Run
pysparkin the prompt—this should automatically launch Jupyter Notebook. - In a new notebook, run this simplified test code (no need for
findsparkif your env variables are correct):from pyspark.sql import SparkSession spark = SparkSession.builder.appName("GlueTest").getOrCreate() df = spark.sql('''select 'spark' as hello ''') df.show()
You should see the expected output:
+-----+ |hello| +-----+ |spark| +-----+
4. Why findspark Wasn’t Working
When you ran just import findspark and findspark.init() with no error, it likely means findspark located your Spark installation—but the underlying Python version mismatch prevented the pyspark import from succeeding. Fixing the Python version resolves this chain issue.
Bonus Tips for AWS Glue Testing
- Ensure your
aws-glue-libsversion matches your Spark version: Glue 2.0 uses Spark 2.4.7, which is close enough to your 2.4.8 setup. - You can add AWS credentials to your environment (via
aws configureor environment variables) if you need to connect to S3 or other AWS services from your local notebook.
内容的提问来源于stack exchange,提问作者DMC

