You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Windows下Python3.14中PySpark调用.show()报Python Worker意外退出错误

Windows 10下PySpark调用DataFrame.show()触发Python Worker崩溃问题

在Windows 10系统中使用PySpark时,调用DataFrame的.show()方法出现错误,任务因Python Worker意外崩溃而失败。

环境信息

  • 操作系统:Windows 10
  • Spark:Apache Spark(PySpark)
  • IDE:VS Code
  • Python版本:3.14

错误表现

调用.show()方法时,Spark任务报错,提示Python Worker意外崩溃,导致任务执行失败(原问题附带3张错误截图)。

复现代码

from pyspark.sql import SparkSession
from pyspark.sql.types import StructType, StructField, StringType

spark = SparkSession.builder \
    .appName("Test") \
    .master("local[*]") \
    .getOrCreate()

emp_schema = StructType([
    StructField("employee_id", StringType(), True),
    StructField("department_id", StringType(), True),
    StructField("name", StringType(), True),
    StructField("age", StringType(), True),
    StructField("gender", StringType(), True),
    StructField("salary", StringType(), True),
    StructField("hire_date", StringType(), True)
])

emp_data = [
    ["001", "101", "John Doe", "30", "Male", "50000", "2015-01-01"],
    ["002", "101", "Jane Smith", "25", "Female", "45000", "2016-02-15"],
    ["003", "102", "Bob Brown", "35", "Male", "55000", "2014-05-01"],
    ["004", "102", "Alice Lee", "28", "Female", "48000", "2017-09-30"],
    ["005", "103", "Jack Chan", "40", "Male", "60000", "2013-04-01"],
    ["006", "103", "Jill Wong", "32", "Female", "52000", "2018-07-01"],
    ["007", "101", "James Johnson", "42", "Male", "70000", "2012-03-15"],
    ["008", "102", "Kate Kim", "29", "Female", "51000", "2019-10-01"],
    ["009", "103", "Tom Tan", "33", "Male", "58000", "2016-06-01"],
    ["010", "104", "Lisa Lee", "27", "Female", "47000", "2018-08-01"],
    ["011", "104", "David Park", "38", "Male", "65000", "2015-11-01"],
    ["012", "105", "Susan Chen", "31", "Female", "54000", "2017-02-15"],
    ["013", "106", "Brian Kim", "45", "Male", "75000", "2011-07-01"],
    ["014", "107", "Emily Lee", "26", "Female", "46000", "2019-01-01"],
    ["015", "106", "Michael Lee", "37", "Male", "63000", "2014-09-30"],
    ["016", "107", "Kelly Zhang", "30", "Female", "49000", "2018-04-01"],
    ["017", "105", "George Wang", "34", "Male", "57000", "2016-03-15"],
    ["018", "104", "Nancy Liu", "29", "Female", "50000", "2017-06-01"],
    ["019", "103", "Steven Chen", "36", "Male", "62000", "2015-08-01"],
    ["020", "102", "Grace Kim", "32", "Female", "53000", "2018-11-01"]
]

# 冗余导入可删除
# from pyspark.sql.types import StructType, StructField, StringType

emp_schema = "employee_id string, department_id string, name string, age string, gender string, salary string, hire_date string"

emp = spark.createDataFrame(emp_data, emp_schema)
emp.show()

解决方案

1. 匹配Python与Spark版本兼容性

Python 3.14属于极新版本,目前Spark 3.5.x及以下版本均未提供官方支持。建议将Python降级到3.8~3.11区间,这是Spark官方明确兼容的版本范围。

2. 配置Windows环境变量

  • 添加PYSPARK_PYTHON环境变量,值为你的Python可执行文件绝对路径(如C:\Python310\python.exe)
  • 添加PYSPARK_DRIVER_PYTHON环境变量,值与上述路径一致,确保Driver和Worker进程使用同一Python解释器
  • 确保Spark安装目录下的bin文件夹已添加到系统PATH中

3. 显式指定SparkSession的Python路径

在创建SparkSession时,直接配置Python路径参数:

spark = SparkSession.builder \
    .appName("Test") \
    .master("local[*]") \
    .config("spark.pyspark.python", "C:/Python310/python.exe") \
    .config("spark.pyspark.driver.python", "C:/Python310/python.exe") \
    .getOrCreate()

4. 排查安全软件拦截

Windows防火墙或杀毒软件可能会拦截Spark的Python Worker进程,可临时关闭相关软件测试是否恢复正常。


内容的提问来源于stack exchange,提问作者Deepika Goyal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.11 16:34:53