You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Jupyter Notebook与PySpark集成调试求助:本地环境配置后运行报错

Fixing PySpark-Jupyter Compatibility Issue (TypeError: an integer is required)

Hey there! Let’s work through this frustrating PySpark + Jupyter problem you’re hitting. That TypeError: an integer is required (got type bytes) error is a classic sign of a Python version mismatch with your Spark installation, and once we fix that, we’ll get your local environment up and running for AWS Glue testing in no time.

Root Cause

You’re using Spark 2.4.8, which was released before Python 3.8 came out. Python 3.8 changed how the types.CodeType constructor works (it now expects integers for certain parameters instead of bytes), and the older cloudpickle library bundled with Spark 2.4.x doesn’t handle this change. That’s exactly what’s causing the stack trace you shared.

Step-by-Step Solutions

1. Create a Python 3.7 Conda Environment (Critical!)

Spark 2.4.x officially supports Python 3.6 and 3.7—skip 3.8+ entirely for this setup. Here’s how to set up a compatible environment:

  • Open Anaconda Prompt and run:
    conda create -n pyspark_glue_env python=3.7
    conda activate pyspark_glue_env
    
  • Install required packages for your workflow:
    conda install jupyter findspark
    pip install aws-glue-libs  # For AWS Glue code testing
    

2. Verify & Update Environment Variables

Double-check these system environment variables (adjust paths to match your installations):

  • SPARK_HOME: Points to your Spark 2.4.8 directory (e.g., C:\spark-2.4.8-bin-hadoop2.7)
  • HADOOP_HOME: Points to your winutils directory (e.g., C:\winutils\hadoop-2.7.1)
  • Add %SPARK_HOME%\bin and %HADOOP_HOME%\bin to your System PATH

For your new Conda environment, set these session-specific variables to ensure PySpark uses the correct Python interpreter:

set PYSPARK_PYTHON=C:\Users\[YourUsername]\anaconda3\envs\pyspark_glue_env\python.exe
set PYSPARK_DRIVER_PYTHON=jupyter
set PYSPARK_DRIVER_PYTHON_OPTS=notebook

3. Test the Fixed Environment

With your Conda environment activated and variables set:

  1. Run pyspark in the prompt—this should automatically launch Jupyter Notebook.
  2. In a new notebook, run this simplified test code (no need for findspark if your env variables are correct):
    from pyspark.sql import SparkSession
    spark = SparkSession.builder.appName("GlueTest").getOrCreate()
    df = spark.sql('''select 'spark' as hello ''')
    df.show()
    

You should see the expected output:

+-----+
|hello|
+-----+
|spark|
+-----+

4. Why findspark Wasn’t Working

When you ran just import findspark and findspark.init() with no error, it likely means findspark located your Spark installation—but the underlying Python version mismatch prevented the pyspark import from succeeding. Fixing the Python version resolves this chain issue.

Bonus Tips for AWS Glue Testing

  • Ensure your aws-glue-libs version matches your Spark version: Glue 2.0 uses Spark 2.4.7, which is close enough to your 2.4.8 setup.
  • You can add AWS credentials to your environment (via aws configure or environment variables) if you need to connect to S3 or other AWS services from your local notebook.

内容的提问来源于stack exchange,提问作者DMC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 05:17:36